Skip to content

Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis

Qijia He, Ruinan Jin, Jun Luo, Shaofeng Zou, Yingbin Liang

cs.LGarXiv:2610.01896

Abstract

Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from O(ε-4) to O(ε-2) as ε0, where 1+ε is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as G-2/5 after tuning the step size, where G is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as G∞, whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.

Create a lesson