Safety Must Survive Self-Improvement: Why Failures Persist and How Agents Recover
Yunbei Zhang, Janet Wang, Saiyue Lyu, Yingqiang Ge, Kaiqu Liang, Zijian Jin, Chandan K Reddy, Jihun Hamm
Abstract
Recursive self-improvement (RSI) allows agents to carry useful changes across generations. Maintaining safety across these generations involves both preventing unsafe behavior from persisting and enabling recovery when failures occur. We study these challenges through a controlled testbed of stateful authorization tasks, where fixed LLM editors optimize executable agent components and independent traces record their effects. Paired interventions separate which revisions pass validation, which program continues running, and which program the editor revises next. After a new authorization dependency invalidates previously tested optimizations, historical scores preserve the same unsafe programs in 22 of 48 framework histories despite a correct alternative in every affected archive. Refreshing scores restores correctness on the original suite, with residual failures on independently composed tests. Failures also persist under an unchanged contract when all proposals are rejected and the failed incumbent remains active. Starting from shared failures, editing the initial implementation instead of the failed one improves recovery, although the advantage varies across editors. Full validation and validated rollback end with fully correct programs in the core trajectory study while retaining over 43% deployment savings. Preserving agent safety requires checking what will run under current conditions and choosing which implementation to edit next.
Create a lesson
Related papers
Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
Zichen Xie, Mrigank Pawagi, Lize Shao et al.
CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
Zheng Fang, Yongmin Li, Yichang Zhang et al.
Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection
Alberick Euraste Djire
Architectural Degradation: How to Measure and to Remediate
Noman Ahmad, Ruoyu Su, Matteo Esposito et al.
Refactoring React Component Hierarchies to Eliminate Prop Drilling
Vangelis Gkinis, Vassilis E. Zafeiris
A Design Theory for AI-Assisted Software Development Derived from Christopher Alexander's Theory of Form
Chien-Tsun Chen, Yu Chin Cheng