RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation
Rongcun Wang, Shi Chen
Abstract
Smart contracts written in Solidity manage assets, permissions, and irreversible state changes, making code generation both useful and security-critical. Repository-level Solidity generation is challenging because models must synthesize complete contracts or libraries while preserving consistency across state variables, modifiers, events, inheritance, external calls, and access-control logic. We present RAPO-Sol, a two-stage training framework for repository-level Solidity code generation. First, Retrieval-Augmented Fine-Tuning (RAFT) augments each training input with similar Solidity examples, helping the model learn recurring contract-level patterns while remaining retrieval-free at inference time. Second, Direct Preference Optimization (DPO) trains the model to prefer reference contracts over close but semantically flawed alternatives. We construct rejected samples using Solidity Semantic-Anchor Perturbation (SAP), which perturbs validation statements, visibility modifiers, data-location keywords, context variables, payment operations, and low-level calls. Experiments on SolidityBench with CodeLlama-7B-Instruct, DeepSeek-Coder-6.7B-Instruct, and Qwen2.5-Coder-7B-Instruct show that RAFT consistently improves over supervised fine-tuning, while SAP-based DPO provides further gains in BLEU and SolidityScore. The full RAFT+DPO pipeline achieves the best performance across all three models, demonstrating complementary benefits from retrieval during training and Solidity-aware preference optimization without adding retrieval cost at inference.
Create a lesson
Related papers
A Case Study in Assuring AI-Written Software
Lindsey Ferris, Sierra Bonilla
Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis
Mandana Ghadamian, David Mohaisen
Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code
Thiago Santos de Moura, Fynn Matuschek, Flavio Toffalini et al.
Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
Anvi Kalpesh Shah, Umamaheswara Sharma B
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
Haibo Jin, Xinjie Li, Peng Kuang et al.
ES-Trace: Auditing Ethical-Sourcing Disclosure of Code Generation Models Beyond Model Cards
Zhuolin Xu, Haibo Wang, Shin Hwei Tan