Skip to content

The Asymptotics of Language Model Alignment with Memory

Haricharan Balasundaram, V. Arvind Rameshwar

cs.CLarXiv:2610.01828

Abstract

Language model (LM) alignment broadly aims to perturb a given LM Q into an aligned LM q such that i) the outputs produced by q and Q are 'close' in probability, ii) q has a higher expected reward than Q. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-n algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an m--length i.i.d. token sequence output by the LM, in the limit as m increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the m--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when m=1 -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.

Create a lesson