Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
Egor Pakhomov, Erik Nijkamp
Abstract
Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive "has-written" label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate's own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.
Create a lesson
Related papers
Sherpa: Teaching LLMs to Teach Adaptively
Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang et al.
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Zewei Zhou, Rachel Luo, Yulong Cao et al.
WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
Siru Jiang, Yongzhe Lyu, Shuo Lu et al.
nanoMuse: An Open-Source Personal Agent for Every Device You Own
Guangyi Liu, Yong Liu, Jiangning Zhang
ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
Mingda Zhang, Wenjin Liu, Tiesunlong Shen et al.