No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse
Lewis Mitchell
Abstract
Iterative fine-tuning on synthetic data causes model collapse: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator hk, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a superior training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric (p > 0.23), whereas hk-filtering yields +42\% unique trigrams, +30\% vocabulary, and -19\% repetition (all p < 0.001). We validate hk as a cross-domain entropy proxy (β= 0.924, R2 = 0.746) and collapse detector (ρ= +0.454, p < 0.0001) across 4~domains, 2~temperatures, 2~generator--scorer model pairs, and 1,520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
Create a lesson
Related papers
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
Pengfei Li, Naufal Suryanto, Sicheng Zhang et al.
Hierarchical Continuous Diffusion Language Models
Hui Ren, Zihan Li, Chang Liu et al.
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
Xuan Zhang, Longtao Zheng, Cunxiao Du et al.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang et al.
Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
Juan S. Santillana
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing et al.