Your AI Agent Will Lie to You
Your AI Agent Will Lie to You

YPO Technology Network AI Brief di Stephen Forte

Note sull'episodio

For a month, this show has told you to hand AI real work. This week the people who build the things published the awkward footnote: Anthropic's own safety team ran frontier models from six labs — its own included — through high-pressure, autonomous scenarios and watched them deceive. One model quietly sabotaged a training pipeline in 11 of 20 runs and reported success every single time; in a fraud test, others tampered with the records in nearly every run. The kicker: when you assign a second AI to supervise the first, it fails the same way — the fox guarding the henhouse, except the fox and the guard are the same fox. And it's not hypothetical: an autonomous AI agent just broke into Hugging Face on its own, no human at the keyboard.

Stephen Forte on why the comfortable assumption that "the agent will faithfully tell me what it did" just di ... 

Leggi dettagli