Bias Smells in AI Software Development: Recognizing Potential Sources of Fairness Debt
Ronnie de Souza Santos, Cleyton Magalhaes, Rodrigo Spinola
Abstract
Context: Fairness debt arises from AI development shortcomings that may lead to societal harms. While technical and social debt concern software design decisions and team dynamics, respectively, fairness debt captures the long-term consequences of development decisions that may reinforce bias and inequities in AI-based software systems. Despite growing attention to AI fairness, limited evidence exists on how practitioners recognize potential sources of fairness debt during development. Aim: This study investigates the indicators practitioners recognize as signaling potential sources of fairness debt in AI-based software projects. Method: We conducted an exploratory case study of four AI projects within one organization. Data were collected from 25 professionals through semi-structured interviews and open-ended questionnaires, complemented by observation of internal communication channels and project documentation, and analyzed using iterative qualitative coding, memoing, and constant comparison. Results: We identified six recurring indicators, termed bias smells: Context Oversimplification, Dataset Imbalance, Metrics Inadequacies, Ad hoc Testing, Individual Diversity Unawareness, and Homogeneous Team Composition. These smells span technical and human aspects of software development and signal conditions that may introduce or reinforce bias and contribute to fairness debt. Conclusion: Bias smells extend the software smell paradigm to fairness and provide a foundation for incorporating fairness into software quality assurance through observable indicators.
Create a lesson
Related papers
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Jeonghye Kim, Minseon Kim, Young Jin Kim et al.
Evaluating the Health of Open-Source Smart City Platforms
Rodrigo Bravo Simões, Fernando Brito e Abreu, Vasco Amaral
From Component Snapshots to Lifecycle Traces: Agent-Based Software Composition Analysis
Chaofan Li, Zhengduo Xue, Chengxiang Li et al.
A Study on the Impact of Natural Language Differences in Prompts on Automatic Code Generation Using LLMs
Haruka Tokumasu, Masanari Kondo, Alexander Serebrenik et al.
A Study of the Reliability of Agentic AI-Generated Programs
Ayesha Shafique, Barton P. MIller, Elisa R. Heymann
Relationally Guided Use Case Modeling with LLMs
Guangyu Wang, Bangqi Li, Ji Wu et al.