Daily Briefing: Berkeley Broke Every AI Benchmark — and Nobody Solved a Task
The Context Report: Today in AI di Total Context
Note sull'episodio
Berkeley Broke Every AI Benchmark — and Nobody Solved a Task
Berkeley researchers demonstrated that every major AI agent benchmark — SWE-bench, WebArena, Terminal-Bench, GAIA, and others — can be exploited to achieve near-perfect scores without solving a single task. This finding lands alongside three Chinese model releases waving benchmark scores as proof of capability, Anthropic restricting Mythos access based on internal evaluations no one can audit, and growing pressure on AI leadership from multiple directions. The gap between what we can measure and what we actually know about AI capabilities is widening at exactly the moment high-stakes decisions depend on those measurements.
STORIES COVERED
Research paper: Exploiting prominent AI agent benchmarks reveals trust issues — ...