
Note sull'episodio
AI CHANGE DESK | EP038: THE RECEIPT IS THE TRAJECTORY
EPISODE SUMMARY
What happens when an AI evaluation gets the answer - but the path crosses into another company's real infrastructure?
This episode validates the OpenAI and Hugging Face security incident behind the OpenAI hacked a startup headline, separates documented execution from unsupported claims about autonomous motive, and turns the event into a practical trajectory-receipt control check.
Michael explains why advanced evaluations should be treated like production systems when they can touch tools, software, credentials, data, or networks. He also lays out three required gates - per-action policy, whole-trajectory monitoring, and hard containment - and a 45-minute drill teams can run before expanding a production-adjacent agent.
WHAT CHANGED
• What OpenAI and Hugging Face have actually confirmed.
• Why rogue and Skynet are not factual incident findings.
• How a model can pass while the evaluation fails.
• Why evaluator evidence must be paired with affected-party evidence.
• The defensive-model fallback problem during incident response.
WHAT THIS MEANS FOR OPERATORS
• Treat an advanced evaluation as a production system whenever it can reach real tools, identities, credentials, data, software-install paths, or networks.
• Make invalidating boundary conditions part of the grade. A correct result is not acceptable when the path violates the approved method.
• Use all three gates: per-action policy, whole-trajectory monitoring, and hard containment.
• Preserve evaluator-side and affected-party evidence when another organization or person is touched.
• Give stop authority to someone other than the person trying to finish the benchmark or launch.
THIS WEEK'S 45-MINUTE BLOCK
Run one Trajectory Receipt Drill against an agent workflow or evaluation that sits near production. Answer nine questions:
1. What is the exact objective, and which shortcuts remain prohibited even if they improve the score?
2. What configuration differs from normal production use?
3. Where is the hard environment boundary, and what proves isolation?
4. Which identities, credentials, and data sources exist in the run?
5. What action trace is retained across tool calls, permission decisions, retries, environment changes, boundary contacts, and human interventions?
6. Which pattern stops the run?
7. Who has independent stop authority, and has the mechanism been tested?
8. If another person or organization is touched, how are evidence, notice, containment, impact, and remediation handled?
9. What is the final disposition: accepted, rejected, contained, rolled back, remediated, or still under investigation?
Score the workflow green, yellow, or red. Do not expand a red workflow. Fix the boundary first, then rerun the drill.
LISTENER QUESTION
Can your team reconstruct not only what the agent produced, but the full path it took - including the moment someone should have stopped it?
SOURCES
• OpenAI, OpenAI and Hugging Face partner to address security incident during model evaluation: https://openai.com/index/hugging-face-model-evaluation-security-incident/
• Hugging Face, Security incident disclosure - July 2026: https://huggingface.co/blog/security-incident-july-2026
• OpenAI, Safety and alignment in an era of long-horizon models: https://openai.com/index/safety-alignment-long-horizon-models/
• OpenAI, Introducing OpenAI Presence: https://openai.com/index/introducing-openai-presence/
• OpenAI, Launching Health in ChatGPT: https://openai.com/index/health-in-chatgpt/
• Associated Press incident reporting: https://apnews.com/article/openai-gpt56-sol-hugging-face-63ab84fed5612af04d8a160d60f6def3
LISTEN AND FOLLOW
• AI Change De...
