Skip to results
MLSift
← Feed
routineOtherEvaluation Replay2608.19269

What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals

Xi Qin, Jizhou Tong

cs.SE cs.AI

Abstract

Benchmarks can run without determining what their results license. We freeze a large evaluation collection and attempt to replay its historical claims. Most units stop because the evidence required for replay is not bound. Where replay is possible, different claims remain stable at different resolutions. We make this otherwise implicit inference step explicit and executable.

Topics

Classified with taxonomy v2 on Sat, 5 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF