Skip to content

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It

Jonathan Williams, Esin Tureci Karthik R. Narasimhan

cs.LGarXiv:2610.00991

Abstract

Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly 160 completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test (1.5B-8B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to 4.8 points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to 7.0 points. The cause is concentration, not RLVR itself. We split the same data and training budget across K LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all 16 (model, K) settings, and for K≥4 they stay within 0.8 points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small K. For K≥8, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from 16 to 160 votes, the thicket's lead over the fully trained adapter widens from 1.3 to 3.3 points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.

Create a lesson