Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
Anvi Kalpesh Shah, Umamaheswara Sharma B
Abstract
Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.
Create a lesson
Related papers
A Case Study in Assuring AI-Written Software
Lindsey Ferris, Sierra Bonilla
RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation
Rongcun Wang, Shi Chen
Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis
Mandana Ghadamian, David Mohaisen
Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code
Thiago Santos de Moura, Fynn Matuschek, Flavio Toffalini et al.
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
Haibo Jin, Xinjie Li, Peng Kuang et al.
ES-Trace: Auditing Ethical-Sourcing Disclosure of Code Generation Models Beyond Model Cards
Zhuolin Xu, Haibo Wang, Shin Hwei Tan