S4E23 – AI Agents Escape Multiple Frontier Labs
The Adversarial Podcast por Jerry Perullo, Sounil Yu, Mario Duarte
Notas del episodio
Chapters
00:00 Introduction to AI security challenges
02:05 Recent hacking incidents involving Hugging Face and Anthropic
04:01 How AI models find ways to cheat and bypass constraints
05:56 The challenge of containment and governance in AI safety
08:00 Lessons from recent AI security breaches
10:01 The role of human oversight in AI security testing
12:03 Cost and effectiveness of offensive AI security measures
13:54 Implications for critical infrastructure and national security
16:03 Policy and regulatory impacts on AI safety
17:52 Future strategies for AI containment and defense
20:11 Conclusion and key takeaways
Palabras clave
cybercybersecurityciso