Skip to content

What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals

Xi Qin, Jizhou Tong

cs.SEarXiv:2608.19269

Abstract

Benchmarks can run without determining what their results license. We freeze a large evaluation collection and attempt to replay its historical claims. Most units stop because the evidence required for replay is not bound. Where replay is possible, different claims remain stable at different resolutions. We make this otherwise implicit inference step explicit and executable.

Create a lesson