Next-token functional estimation
Milind Nakul, Vidya Muthukumar, Ashwin Pananjady
Abstract
Suppose we observe the first n points of a sequence of random variables having length n+1, and wish to estimate a functional of the unobserved final point and the empirical measure of the n observed training points. Such next-token functionals include the probability that the next token is novel (also known as the surprise probability), the tail probability of the minimum distance between the next token and training points, and the test error of a classifier trained on the observed points. All of these quantities are classically estimated by the leave-one-out method, which is inconsistent under temporal dependence. We propose a leave-a-window-out estimator, which deletes a window of length τ after each index before forming the empirical measure and reduces to leave-one-out at τ= 1. Under natural assumptions, we show that the error of our estimator decays at a parametric rate for any stationary β-mixing process that also admits a Marton coupling. Our results thus cover several natural functionals on a large class of stochastic processes. We complement these upper bounds with a sharp minimax lower bound for estimating the surprise probability on mixing Markov chains. Simulations on Markov chains, moving-average processes, and autoregressive processes show that our estimator succeeds in many scenarios where leave-one-out and add-constant baselines fail.
Create a lesson
Related papers
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
Sho Kawano, Zehang Richard Li, Paul A. Parker
TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models
Jingbo Liu, Zhiyuan Yu
Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs
Zhenlin Yao, Wei Xiong
Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning
Weiwei Wang, Yuqiang Li, Xianyi Wu et al.
Error bounds in Sobolev norms for approximations with norm constrained ReLU neural networks
Xianjun Li, Yunfei Yang
Null importance: Disentangling relevance for interpretable machine learning
Garvesh Raskutti, Kris Sankaran, Jiaxin Ye