Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses
Xinke Jiang, Zhixin Zhang, Zhibang Yang, Jiaran Gao, Rihong Qiu, Shijin Chen, Xu Chu, Junfeng Zhao, Yasha Wang
Abstract
Large language model agents increasingly solve long-horizon tasks through multi-agent harnesses in which a central agent coordinates specialized sub-agents, tools, and environments. Training the central policy in such a harness raises two challenges. First, an action label is a low-cardinality decision, whereas its args form a high-dimensional conditional sequence; optimizing both with a shared sequence-level signal can produce conflicting gradients. Second, dynamic scheduling creates interdependent sessions with branches, parallel calls, and rewritten contexts, which cannot be faithfully reduced to one flat token sequence. We introduce Harness-RL, a structured reinforcement learning framework that combines Conflict-Aware Policy Optimization (CAPO) with interface-level black-box trajectory construction. The black-box component captures Interface Call Records, builds per-session prefix trees, and aligns outcome and process rewards with trainable tokens. CAPO uses forward activations to identify parameter partitions associated with action and args tokens, then routes their policy gradients to the corresponding subspaces. Harness-RL supports both central-only and joint multi-agent training. Across seven multi-hop question answering and agentic retrieval benchmarks, it reaches average F1 scores of 42.93 and 47.79 with Qwen2.5-1.5B and Qwen2.5-3B, respectively, while ablations validate the contribution of CAPO and favor central-only optimization in the evaluated setting. Our code is available at https://github.com/jiangxinke/Harness-RL.
Create a lesson
Related papers
AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
Xinke Jiang, Yue Fang, Zhibang Yang et al.
Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps
Sagar Srinivas Sakhinana, Venkataramana Runkana
Generalizable Multi-Agent Planning from Signal Temporal Logic Specifications via Diffusion
Joe Eappen, Zikang Xiong, Shreyash S. Iyengar et al.
ASTRA - Agentic System for Ticket Resolution and Analysis
Shashidhar Reddy Javaji, Mohamed Trabelsi, Jin Cao et al.
One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles
Zhichen Zeng, Huiyuan Chen, Jingru Cheng et al.
Dynamic Haven Selection for Multi-Agent Pickup and Delivery in Constrained Warehouses
Taisei Hirayama, Kohei Yoshida, Hiroki Sakaji et al.