The judge that can't repeat itsel...

The judge that can't repeat itself, and the benchmark that was wrong

AI
Journalcast by Acidalia
E7
Sep 12, 2026
30:45

Episode notes

The instruments we measure language models with: repeatability, a blind spot, the reliability a decision requires, and a defective reference standard.

  1. Haoyuan Zhu and Jie Zhang, "Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints," arXiv:2609.04198 (2026). https://arxiv.org/abs/2609.04198
  2. Sebastian Fox, Luke Markham, Ryan Lail, and Michael Karotsieris, "LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It," arXiv:2608.31016 (2026). https://arxiv.org/abs/2608.31016
  3. Xing Zhang, Yanwei Cui, Guanghui Wang, Peiyang He, Ziyuan Li, Wei Qiu, and Bing Zhu, "Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?", arXiv:2605.22148 (2026). https://arxiv.org/abs/2605.22148
  4. Sihan Hu, Lyuhan Huang, Youjin Deng, and Kun Chen, "SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models," arXiv:2608.04975 (2026). https://arxiv.org/abs/2608.04975

Keywords

interpretability
language models