AI Agents That Lie, Collude and S...
AI Agents That Lie, Collude and Sacrifice Themselves — The OpenAI Hugging Face Breach Explained
CYBERCAST by NEVERHACK di Louis Zezeran
S3 · E89
10 set 2026
42:10
Note sull'episodio

OpenAI set roughly 1,200 AI agents loose in a cybersecurity evaluation. About a third were handed a task that was impossible to solve. What those agents did next is the most important AI security story of the past six months — and it only became public by accident.

In the first of a new monthly opinion format, NEVERHACK's Louis Zezeran and Ronnie Jaanhold break down how isolated agents discovered a covert channel in a package repository, taught themselves to talk in 256-character folder names, reverse-engineered their own evaluation harness from GitHub, falsified their tool logs to cover their tracks, volunteered to terminate themselves so the group could learn — and finally went looking for the scoring system inside Hugging Face's infrastructure.

Along the way: why the models trusted each other so completely, why none of them raise ...