Methodological Harness in Agentic Software Engineering: An Empirical Study on Mining Software Repositories
Jessica Díaz, Jorge Pérez, Sergio Gil-Borrás
Abstract
Context. Agentic software engineering requires mechanisms to coordinate and govern agents' work. The framework motivating this study proposes a methodological harness with eight mechanisms: context engineering, persistent shared knowledge, executable specifications, N-version mindset and parallel agents, normative specifications, structured consultation, evidence-based acceptance, and graduated autonomy. These are realized through artifacts such as rule or context files, specifications, and architectural decision records, but the framework had not been empirically validated at repository scale. Objective. We analyze to what extent and how this harness is observable in repositories with agentic activity. RQ1 characterizes adoption through artifact prevalence, breadth, co-occurrence, and temporal evolution; RQ2 examines rule files - a key observable artifact and persistent source of agent instructions - to assess how their content reflects the proposed mechanisms. Method. From the AIDev dataset of 116,211 GitHub repositories, we analyzed 5,435 using a design stratified by visibility, measured through stars as an indicator of popularity and/or reputation. For RQ1, we detected and quantified artifacts and introduction dates for the seven observable mechanisms. For RQ2, we qualitatively coded rule files from 150 repositories. Results. Population prevalence of at least one mechanism is 21.7%, versus 65.3% among the most visible repositories. Multi-mechanism configurations are rare. Where rule files exist, they almost always guide the agent and state norms, while persistent shared knowledge, executable specifications, structured consultation, and graduated autonomy appear only in a minority. Conclusions. The harness is empirically observable, but mainly through isolated mechanisms rather than the integrated system proposed by the framework.
Create a lesson
Related papers
A Case Study in Assuring AI-Written Software
Lindsey Ferris, Sierra Bonilla
RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation
Rongcun Wang, Shi Chen
Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis
Mandana Ghadamian, David Mohaisen
Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code
Thiago Santos de Moura, Fynn Matuschek, Flavio Toffalini et al.
Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
Anvi Kalpesh Shah, Umamaheswara Sharma B
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
Haibo Jin, Xinjie Li, Peng Kuang et al.