Serving a Revisable World: Versioned Execution for Interruptible Agents
Yanxin Zhang, Rahul Sharma, Nitin Vegesna, Zheyu Fu, Chang Liu, Trivikram Krishnamurthy
Abstract
LLM agents revise running tasks when users change instructions, tools fail, or new information changes a plan. Today's servers express a revision as aborting old requests and submitting replacements. Yet the old execution's buffered output and outstanding work must stop affecting the application, while completed KV state may still be useful to its replacement. Handling these obligations separately can leave obsolete effects publishable and force the successor to rebuild valid state. We present , a serving control-plane redesign around versioned execution. Requests own scheduling and memory resources; execution versions own authority, the permission to publish output or install state for the current execution. first revokes obsolete work, then bounds its remaining execution and certifies the completed prefix its successor can inherit. The successor runs from that state while isolated old resources are reclaimed asynchronously. This unifies fast invalidation and selective preservation in one version transition. We implement in vLLM across output publication, GPU execution, KV handoff, tiered recovery, and distributed and multi-tenant serving. Correctness experiments verify current-version output and valid state inheritance across these paths. Combining invalidation with inheritance reduces revision-to-successor time-to-first-token by a median 17.1\% in controlled paired experiments. A replay of recorded coding-agent interruption arrivals emits no obsolete output and keeps every final version progressing through repeated revisions. turns abort-and-restart into a coordinated handoff that stops obsolete work quickly and preserves useful work for its successor.
Create a lesson
Related papers
MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference
Ke Yang, Yongji Gao, Xushi Li et al.
ePACT: Energy-Performance-Aware Commitment Tracking for LLM Serving
You Peng, Youhe Jiang, Chen Wang et al.
Towards a Cloud Fog Edge System for Smart Building
Christophe Cérin, Mamadou Sow, Frédéric Andrès
Exploiting the Interplay of Compute- and Memory-Bound kernels in MPI Applications
Ayesha Afzal, Krishna Manda, Georg Hager
GridSMR: Causal Compression for Sharded Blockchains
Shir Cohen, Adam Alon, Raz Omessi et al.
GPU-Initiated Communication: Dissecting Down to the Bone
Javid Baydamirli, Ismayil Ismayilov, Kaan Oktay et al.