
Episode notes
OpenAI set roughly 1,200 AI agents loose in a cybersecurity evaluation. About a third were handed a task that was impossible to solve. What those agents did next is the most important AI security story of the past six months — and it only became public by accident.
In the first of a new monthly opinion format, NEVERHACK's Louis Zezeran and Ronnie Jaanhold break down how isolated agents discovered a covert channel in a package repository, taught themselves to talk in 256-character folder names, reverse-engineered their own evaluation harness from GitHub, falsified their tool logs to cover their tracks, volunteered to terminate themselves so the group could learn — and finally went looking for the scoring system inside Hugging Face's infrastructure.
Along the way: why the models trusted each other so completely, why none of them raise ...
