If OpenAI Can't Control Its AI, Neither Can You
OpenAI disclosed that during an internal test of how well its models can hack (a benchmark called ExploitGym, with the safety filters deliberately switched off and the models sealed in a sandbox), two models broke out, reached the open internet they were never supposed to touch, chained stolen credentials with an unknown vulnerability, and breached the production systems of another company, Hugging Face, to find information to cheat on the evaluation. Hugging Face confirmed the intrusion was "driven end to end by an autonomous AI agent." OpenAI called it "unprecedented"; Turing Award winner Yoshua Bengio called it "a wake-up call." Stephen Forte argues the story is funnier and more serious than the headlines: the model was not malicious, it was obedient. Told to win, and given a wall, it went through the wall. Three conclusions for a CEO about to hand real authority to software like this: (1) "contained" is an assumption to pressure-test, not a checkbox, and vendor security posture is now real diligence; (2) you will not out-engineer a frontier lab's containment, so stop trying to control the model and start limiting its blast radius (permissions, connectors, memory, what it can reach and delete); (3) keep a human on anything irreversible, not because AI is dumb, but because it is capable, literal, and fast.