ES-Trace: Auditing Ethical-Sourcing Disclosure of Code Generation Models Beyond Model Cards
Zhuolin Xu, Haibo Wang, Shin Hwei Tan
Abstract
Code generation models have been increasingly used in software development, but their development raises ethical-sourcing concerns involving intellectual property, privacy, fairness, labour practices, and environmental impact. Although prior work has defined ethical-sourcing criteria for code generation, it remains unclear how much evidence existing models disclose and where that evidence can be found. We introduce ES-Trace, a framework for ethical-sourcing disclosure audits that traces disclosed evidence beyond model cards using the Model Documentation Traceability Graph (MDTG), which represents relationships among models, versions, and documentation artifacts. We apply ES-Trace to 26 models from 10 publishers across 77 documents and 20 ES-CodeGen aspects. Model-card-only auditing yields a mean score of 1.77/5, while resolving the declared references increases it to 2.82/5, with most of the increase arising from documents that the publisher declares in structured metadata. The key findings of our study include: (1) social and labour-related aspects remain poorly documented, even when expanding the audit to the full documentation scope, (2) resolving documentation references substantially increases observed disclosure, raising the mean score from 1.77/5 to 2.82/5, (3) model-card-only audits can mischaracterize release-level disclosure changes, and (4) documentation mismatches can associate evidence with the wrong model or version, highlighting the need for explicit model--version binding and consistency across documentation artifacts. Our study calls for reference-aware ethical-sourcing disclosure audits, explicit model--version binding and consistency across documentation artifacts, and stronger documentation of currently underreported social and labour-related aspects.
Create a lesson
Related papers
A Case Study in Assuring AI-Written Software
Lindsey Ferris, Sierra Bonilla
RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation
Rongcun Wang, Shi Chen
Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis
Mandana Ghadamian, David Mohaisen
Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code
Thiago Santos de Moura, Fynn Matuschek, Flavio Toffalini et al.
Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
Anvi Kalpesh Shah, Umamaheswara Sharma B
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
Haibo Jin, Xinjie Li, Peng Kuang et al.