

The judge that can't repeat itself, and the benchmark that was wrong
AI
Journalcast by Acidalia
E7
Sep 12, 2026
30:45
Episode notes
The instruments we measure language models with: repeatability, a blind spot, the reliability a decision requires, and a defective reference standard.
- Haoyuan Zhu and Jie Zhang, "Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints," arXiv:2609.04198 (2026). https://arxiv.org/abs/2609.04198
- Sebastian Fox, Luke Markham, Ryan Lail, and Michael Karotsieris, "LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It," arXiv:2608.31016 (2026). https://arxiv.org/abs/2608.31016
- Xing Zhang, Yanwei Cui, Guanghui Wang, Peiyang He, Ziyuan Li, Wei Qiu, and Bing Zhu, "Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?", arXiv:2605.22148 (2026). https://arxiv.org/abs/2605.22148
- Sihan Hu, Lyuhan Huang, Youjin Deng, and Kun Chen, "SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models," arXiv:2608.04975 (2026). https://arxiv.org/abs/2608.04975
Keywords
interpretability
language models