OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
Babak Barazandeh, Connor Swanson, Chinmay Kulkarni, Nikhil Mungel
Abstract
Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior or evaluating logs post-hoc. The first adds cost and latency to every step; the second delivers its verdict after the run, when tokens are burned and damage is done. To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step. We study this problem in three regimes of decreasing access: full reference access (historical runs and tool schemas), intermediate access (only tool schemas), and no prior knowledge (only step logs as generated). Expectation of OnTrack's monitoring capabilities reduces as data access drops, ranging from plan violation detection to identifying loops, stalls, and repeated tool calls. Finally, we evaluate OnTrack using SWE-bench trajectories. Based on the first 8 steps, our method ranks failing trajectories below succeeding ones better than content similarity approaches (+0.057 AUROC). With an abort policy, we save about 18% of compute that would be burned on failing runs, where 83% of interrupted runs were actually heading to failure (5 out of 6 aborts were correct).
Create a lesson
Related papers
On the estimation and validity of AI time horizons---a statistical look at the METR plot
Drew T. Nguyen, William Fithian
BrickBench: Evaluating Agentic Brick Design
Peter Kulits, Yiqing Xu, R. Kenny Jones et al.
Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
Erin Crawley, Hidenori Tanaka
Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
Christopher M. Stewart, Preston Botter, Natalie Sarabosing et al.
HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
Qun Dai, Liangjian Wen, Jiang Duan et al.
GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving
Jialu Wang, Ruichen Zhang, Xiaoou Liu et al.