Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It
Jonathan Williams, Esin Tureci Karthik R. Narasimhan
Abstract
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly 160 completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test (1.5B-8B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to 4.8 points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to 7.0 points. The cause is concentration, not RLVR itself. We split the same data and training budget across K LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all 16 (model, K) settings, and for K≥4 they stay within 0.8 points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small K. For K≥8, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from 16 to 160 votes, the thicket's lead over the fully trained adapter widens from 1.3 to 3.3 points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
Create a lesson
Related papers
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Jichao Jiang, Cristian McGee, El Houcine Bergou et al.
FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
Akshay Balsubramani
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
Shuo Xing, Zilin Dai, Chengyuan Qian et al.
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Cristian McGee, El Houcine Bergou, Aritra Dutta
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.