Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Cullen Anderson, Narmeen Fatimah Oozeer, Jeff M. Phillips
Abstract
Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can exceed chance without performing the intended reasoning. We show that this failure mode is detectable and can be exploited by downstream classifiers. In TruthfulQA, a simple six-feature logistic classifier achieves substantial accuracy in separating correct from incorrect answers. We further show that similar surface-level artifacts are present in additional benchmarks. To counteract this, we developed a general mechanism to clean them by removing the most leakage-reinforcing pairs. We release a version of TruthfulQA with surface-feature leakage reduced close to chance and provide a mechanism, Audit-Prune, so that the datasets can be cleaned before release.
Create a lesson
Related papers
Type Diversity Enables Transformers to Generalise Compositionally
Anssi Moisio, Mathias Creutz, Mikko Kurimo
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Zhiwei Li, Lei Zhu, Hao Gu et al.
Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents
Yunqi Lu, Tyler Baumgartner, Nikhil Johri et al.
Expert-Space Exploration in MoE Reinforcement Learning
Hongyi He, Zhenghao Lin, Xiao Liu et al.
Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning
Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani et al.
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie et al.