Skip to content

The Variance of Thought: Policy Variance, Critical Forks, and Local Credit Assignment

Yingru Li

cs.LGarXiv:2608.22467

Abstract

Long-horizon language-model tasks --- multi-step reasoning and tool-using agents alike --- are limited by credit assignment. We analyze it through the policy variance σπ2(s)=Varaπ[Qπ(s,a)], which in a deterministic MDP is the sole source of return variance and is injected in discrete pulses at states we call critical forks. Three results follow. (i) Policy variance is a discovery budget: observing an action of advantage c requires Ω(c2/σπ2(s)) draws, a bound that is exact on the canonical two-point fork. (ii) Policy variance is bounded by the policy's Gini dispersion, σπ2(s) 1-\|π(·|s)\|22, a rollout-free necessary condition for criticality computable from logits alone. (iii) The remaining horizon sets the estimation cost: at a fork whose downstream success probability is P, the Monte Carlo advantage estimate has signal-to-noise ratio of order P, so its sample cost scales as 1/P --- a cost that branched sampling shares. Bootstrapping removes it by converting a product of survival probabilities into a sum, provided the value representation is multiplicatively accurate, which argues for log-value parameterization.

Create a lesson