What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals
Xi Qin, Jizhou Tong
cs.SE cs.AI
Abstract
Benchmarks can run without determining what their results license. We freeze a large evaluation collection and attempt to replay its historical claims. Most units stop because the evidence required for replay is not bound. Where replay is possible, different claims remain stable at different resolutions. We make this otherwise implicit inference step explicit and executable.
Topics
Classified with taxonomy v2 on Sat, 5 Sept 2026.