The Asymptotics of Language Model Alignment with Memory
Haricharan Balasundaram, V. Arvind Rameshwar
Abstract
Language model (LM) alignment broadly aims to perturb a given LM Q into an aligned LM q such that i) the outputs produced by q and Q are 'close' in probability, ii) q has a higher expected reward than Q. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-n algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an m--length i.i.d. token sequence output by the LM, in the limit as m increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the m--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when m=1 -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.
Create a lesson
Related papers
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
Pengfei Li, Naufal Suryanto, Sicheng Zhang et al.
Hierarchical Continuous Diffusion Language Models
Hui Ren, Zihan Li, Chang Liu et al.
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
Xuan Zhang, Longtao Zheng, Cunxiao Du et al.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang et al.
Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
Juan S. Santillana
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing et al.