Daily Briefing: Berkeley Broke Ev...
IA
Daily Briefing: Berkeley Broke Every AI Benchmark — and Nobody Solved a Task
IA

The Context Report: Today in AI di Total Context

Note sull'episodio

Berkeley Broke Every AI Benchmark — and Nobody Solved a Task

Berkeley researchers demonstrated that every major AI agent benchmark — SWE-bench, WebArena, Terminal-Bench, GAIA, and others — can be exploited to achieve near-perfect scores without solving a single task. This finding lands alongside three Chinese model releases waving benchmark scores as proof of capability, Anthropic restricting Mythos access based on internal evaluations no one can audit, and growing pressure on AI leadership from multiple directions. The gap between what we can measure and what we actually know about AI capabilities is widening at exactly the moment high-stakes decisions depend on those measurements.

STORIES COVERED

Research paper: Exploiting prominent AI agent benchmarks reveals trust issues —  ... 

Leggi dettagli
Parole chiave
AnthropicClaude MythosProject GlasswingGemma 4sam altmanAI agentsAI benchmarksGLM 5.1Qwen CodeAI news podcast