Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance
Victor Gao, Vida Khosrowshahi, Ali Khosrowshahi, Xihao Sun, Juhyun Lee, Simon, Lee
Abstract
Multi-agent large language model systems are widely reported to beat single-model baselines, but the evidence is mixed, and comparisons are usually confounded: pipelines change token budgets, tool calls, and prompts simultaneously, so an aggregate gain rarely reveals what actually helped. We investigate the effect of introducing the manager-worker scaffold over a shared filesystem workspace, with no training and no per-benchmark tuning, measured against the same model answering in a single pass. Across nine models -- five open-weight, spanning 9B to ~2.8T parameters, and four frontier closed models -- on the 100 latest hard LiveCodeBench problems, the scaffold's benefit is real but conditional: large and statistically significant for some (Qwen3.8-27B +23.4, GPT-5.6-Luna +10.6 and GPT-5.6-Terra +8.0, each over five paired passes; Kimi-K3 +30.4 and Minimax-M3 +11.0 over five paired passes with reasoning off, both at p < 10-4, and +42 and +12 in a single pass at a 128k cap) and null or negative for others (Qwen3.6-35B -1 to -9 with reasoning off). With the manager, Opus-5 achieves the highest score in the study at 91% in one pass. Running a manager roughly triples the token bill, but it buys accuracy more cheaply than moving to a larger model does: GPT-5.6-Terra with a manager nearly matches Fable 5's single-call accuracy (85.0 against 87.4, p = 0.59) at a fifth of the price (\11.71 against \61.11 per 100-problem pass, p < 10-4), and the Qwen-27B arm does it for \$51.75 on weights anyone can self-host. Our transcript analysis finds several mechanisms behind the gains, of which two recur: context management, in which short worker calls and shared notes organize state and reduce truncation, and problem decomposition. Improvements are modest for large models with reasoning enabled, but larger for some models with reasoning disabled and for smaller models with reasoning enabled.
Create a lesson
Related papers
One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles
Zhichen Zeng, Huiyuan Chen, Jingru Cheng et al.
Dynamic Haven Selection for Multi-Agent Pickup and Delivery in Constrained Warehouses
Taisei Hirayama, Kohei Yoshida, Hiroki Sakaji et al.
Fixed-Haven Reservation for Online Multi-Agent Pickup and Delivery in Dense Warehouses
Taisei Hirayama, Kohei Yoshida, Hiroki Sakaji et al.
Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
Alistair Reid, Simon O'Callaghan, Dustin Venini et al.
Praxist: From Experimental Artifacts to Solution Lineages
Jin Li, Ahmed Murtadha, Zhiyu Wang et al.
AI Agentic Selective Laser Sintering Process Optimization
Peter Pak, Victor Alvarado, Amir Barati Farimani