Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection
Alberick Euraste Djire
Abstract
Code detectors can become obsolete as code-generating models evolve: a detector validated on one generation of models may not transfer to the next. We call this limited useful life a detector half-life. We evaluate eight general-purpose LLM judges and three dedicated detectors on human-written code and code produced by seven generators across C++, Java, and Python. Our results reveal two problems. First, performance varies considerably across generators and prompting strategies, suggesting that some detectors rely on generator-specific patterns rather than general evidence of code provenance. Second, accuracy can conceal severe prediction bias. DetectCodeGPT and GPT-Sniffer achieved an accuracy of 0.50 but an F1 score of 0.00 across all generators because they classified almost every sample as AI-generated. However, general-purpose LLM judges achieved stronger accuracy and F1 scores. Our results show that general-purpose LLMs are promising training-free judges of code provenance and can outperform dedicated detectors. However, their reliability depends on the judge model, the code generator, and the prompting strategy. We therefore recommend evaluating LLM judges across multiple generators and reporting macro-F1 alongside class-specific precision and recall.
Create a lesson
Related papers
Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
Zichen Xie, Mrigank Pawagi, Lize Shao et al.
CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
Zheng Fang, Yongmin Li, Yichang Zhang et al.
Architectural Degradation: How to Measure and to Remediate
Noman Ahmad, Ruoyu Su, Matteo Esposito et al.
Refactoring React Component Hierarchies to Eliminate Prop Drilling
Vangelis Gkinis, Vassilis E. Zafeiris
A Design Theory for AI-Assisted Software Development Derived from Christopher Alexander's Theory of Form
Chien-Tsun Chen, Yu Chin Cheng
CompProv Produces Machine Readable Graphs Encoding Microscopic Algebraic Provenance for Reproducible Computation
Minas Abramyan, Mohammed Alaa Ala'anzy, Nasir Saeed