AI Change Desk

AI Change Desk

por Michael Hanna-Butros Meyering

AI Change Desk | EP017: Merchant Control Check

AI shopping is getting more complicated in a way that looks neat in demos and messy in operations. This episode follows EP014 and asks the tighter version of the same question: once discovery starts in ChatGPT, Google AI Mode, or another AI shopping surface, who actually owns the sale, the attribution, the checkout path, and the support policy that comes after it? Why OpenAI’s shift toward product discovery and merchant-controlled checkout matters Why Shopify’s agentic storefront tools make AI shopping feel more like channel ops than hype Why Google’s personalization and protocol work make QA and merchandising harder to reproduce Why “we showed up in the answer” is still not a sufficient success metric OpenAI: https://openai.com/index/powering-product-discovery-in-chatgpt/ OpenAI Help: https://help.openai.com/en/articles/11128490-shopping-with-chatgpt-search Shopify: https://www.shopify.com/news/agentic-commerce-momentum Shopify Help: https://help.shopify.com/en/manual/online-sales-channels/agentic-storefronts/chatgpt Google: https://blog.google/products-and-platforms/products/search/personal-intelligence-expansion/ Google India: https://blog.google/intl/en-in/products/explore-communicate/new-ways-google-is-using-ai-to-make-shopping-easier/ Google UCP updates: https://blog.google/products-and-platforms/products/shopping/ucp-updates/ Search Engine Land: https://searchengineland.com/google-updates-universal-commerce-protocol-to-help-retailers-sell-on-the-open-agentic-web-456891 EP017 Practitioner Worksheet — AI Commerce Control Check EP014: Commerce Surface Check EP015: Retained Artifact Check EP016: National Capacity Check

AI Change Desk | EP013: Career Infrastructure Check

Summary AI is becoming career infrastructure before most schools, employers, and training systems know how to teach it, measure it, or distribute its benefits evenly. This episode looks at the education capability gap, worker compensation behavior, and the institutional response now forming around AI-shaped work. What changed OpenAI argues that education systems need to close an AI capability gap as college-age adults become the biggest adopter cohort and advanced student users still lag well behind power-user behavior. OpenAI says Americans are sending nearly 3 million messages per day to ChatGPT about wages, compensation, or earnings, making AI a live part of worker pay and career decisions. Microsoft launched Elevate for Educators and free student career subscriptions with Copilot features, showing a two-track response: train the teacher and equip the student. Microsoft and Victoria University launched a Datacentre Academy, signaling that AI-driven infrastructure demand is already reshaping workforce pipelines and training priorities. What this means Access is not the same as readiness. Fluency is not the same as frequent use. Institutions now have to answer career questions with more specificity, speed, and trust than they did before AI became the default guide in the browser. Action block — Career infrastructure sweep (45 minutes) Pick one career-facing workflow: internship prep, internal mobility, salary benchmarking, or educator training. Identify where people are already using AI in that workflow. Find one place where AI is faster than your official guidance. Add one verification step and one named owner. Define what “good use” looks like in plain language. Sources OpenAI: Ensuring AI use in education leads to opportunity https://openai.com/index/ai-education-opportunity/ OpenAI: Equipping workers with insights about compensation https://openai.com/index/equipping-workers-with-insights-about-compensation/ Microsoft: Elevate for Educators and new AI-powered tools https://news.microsoft.com/source/2026/01/15/microsoft-expands-its-commitment-to-education-with-elevate-for-educators-program-and-new-ai-powered-tools/ Microsoft and Victoria University: Datacentre Academy https://news.microsoft.com/source/asia/2026/03/27/datacentre-academy-vu/ Disclosure AI-assisted tools were used in parts of research and production support. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and not representative of any organization.

AI Change Desk | EP012: Work Visibility Check

Bonus
AI is moving from side-chat into the live work surface. That means the next management problem is not just launch. It is visibility. Can you tell where adoption is real, where it is helping, and where the rollout is mostly theater? This episode covers: write actions moving AI deeper into connected Google and Microsoft apps, OpenAI's workspace analytics, analytics viewer role, and impact-survey layer, and one practical Adoption Visibility Sweep you can run before Friday. OpenAI's March 13 enterprise release notes show ChatGPT supporting write actions for connected Google Docs, Google Sheets, and calendar apps, plus Microsoft Outlook email and calendar actions. OpenAI's workspace analytics rollout includes an analytics viewer role, and the March 20 release notes added Admin-created surveys and moved OpenAI-created impact surveys to begin on or after March 31. OpenAI's March 5 Adoption news channel makes the vendor shift clear: adoption visibility is now part of the product story. OpenAI's March 11 Wayfair case study gives a concrete example of workflow-level deployment with measurable, vendor-reported results. If AI is now editing the work where the work already lives, leaders need a cleaner way to tell: whether usage is real, whether outcomes improved, and where friction is still hiding. Run a 45-minute Adoption Visibility Sweep: Pick one workflow. Name the artifact that matters. Track usage, outcome, and friction. Ask one manager where the change is real and where it is still cosmetic. Make one Friday decision: train, simplify, standardize, or pause. OpenAI Help Center release notes: https://help.openai.com/en/articles/10128477-chatgpt-enterprise-edu-release-notes OpenAI workspace analytics: https://help.openai.com/en/articles/10875114-workspace-analytics-for-chatgpt-enterprise-and-edu OpenAI adoption news channel: https://openai.com/index/introducing-the-adoption-news-channel/ OpenAI x Wayfair case study: https://openai.com/index/wayfair/ Microsoft Wave 3: https://www.microsoft.com/en-us/microsoft-365/blog/2026/03/09/powering-frontier-transformation-with-copilot-and-agents/ AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization.

AI Change Desk | EP011: Control Surface Check

AI CHANGE DESK | EP011: CONTROL SURFACE CHECK EPISODE SUMMARY AI is getting harder to manage for one simple reason: it is disappearing into the normal work surface. This week’s episode looks at three connected signals: • Google pushing Gemini deeper into Docs, Sheets, Slides, and Drive • Anthropic research showing people question polished AI output less once it looks finished • OpenAI positioning GPT-5.4 for professional work, which turns model choice into a cost, confidence, and review-burden decision If episode 8 was validate before you scale, episode 9 was harden the controls, and episode 10 was name the owners at the handoff, episode 11 is the next layer: what happens when AI stops feeling like a separate tool and starts feeling like ordinary work. WHAT CHANGED • Google is embedding Gemini more deeply into the files people already live in, making AI feel less like a separate stop and more like part of the default work surface. • Anthropic’s AI Fluency Index found that users iterate a lot, but they become less critical once Claude produces polished artifacts like code, documents, and interactive outputs. • OpenAI is positioning GPT-5.4 for professional work and saying it improves factual performance versus GPT-5.2, which makes model choice less about hype and more about acceptable error and review burden. WHAT THIS MEANS FOR OPERATORS The management problem is no longer just tool approval. It is: • where inside normal work the human still needs to slow down • how teams keep skepticism alive after output starts looking finished • which workflows deserve the fastest model versus the most trusted model WHAT I’D DECIDE BY FRIDAY 1. Pick one default work surface and mark three friction points where a human has to slow down. 2. Teach one collaboration habit people will actually use: what is missing, what should I verify, or where is confidence weak. 3. Separate the fast model from the trusted model instead of pretending one default fits every workflow. LISTENER QUESTION Where is your bigger gap right now: noticing AI inside normal work, challenging polished output, or choosing the right model for the job? LISTEN AND WATCH • Episode page: https://michaelhbm.com/AiChangeDesk/episodes/ep011-control-surface-check • Archive: https://michaelhbm.com/AiChangeDesk • Apple Podcasts: https://podcasts.apple.com/us/podcast/ai-change-desk/id1876677295 • Spotify: https://open.spotify.com/show/5X1sLLTeULqFCdt7aaisGD SOURCES • https://openai.com/index/introducing-gpt-5-4/ • https://openai.com/index/introducing-gpt-5-2/ • https://techcrunch.com/2026/03/05/openai-launches-gpt-5-4-with-pro-and-thinking-versions/ • https://blog.google/products-and-platforms/products/workspace/gemini-workspace-updates-march-2026/ • https://techcrunch.com/2026/03/10/google-rolls-out-new-gemini-capabilities-to-docs-sheets-slides-and-drive/ • https://www.anthropic.com/research/AI-fluency-index • https://www.forbes.com/sites/danfitzpatrick/2026/02/23/anthropics-new-ai-index-shows-what-sets-top-ai-users-apart/

AI Change Desk | EP038: The Receipt Is the Trajectory

AI CHANGE DESK | EP038: THE RECEIPT IS THE TRAJECTORY EPISODE SUMMARY What happens when an AI evaluation gets the answer - but the path crosses into another company's real infrastructure? This episode validates the OpenAI and Hugging Face security incident behind the OpenAI hacked a startup headline, separates documented execution from unsupported claims about autonomous motive, and turns the event into a practical trajectory-receipt control check. Michael explains why advanced evaluations should be treated like production systems when they can touch tools, software, credentials, data, or networks. He also lays out three required gates - per-action policy, whole-trajectory monitoring, and hard containment - and a 45-minute drill teams can run before expanding a production-adjacent agent. WHAT CHANGED • What OpenAI and Hugging Face have actually confirmed. • Why rogue and Skynet are not factual incident findings. • How a model can pass while the evaluation fails. • Why evaluator evidence must be paired with affected-party evidence. • The defensive-model fallback problem during incident response. WHAT THIS MEANS FOR OPERATORS • Treat an advanced evaluation as a production system whenever it can reach real tools, identities, credentials, data, software-install paths, or networks. • Make invalidating boundary conditions part of the grade. A correct result is not acceptable when the path violates the approved method. • Use all three gates: per-action policy, whole-trajectory monitoring, and hard containment. • Preserve evaluator-side and affected-party evidence when another organization or person is touched. • Give stop authority to someone other than the person trying to finish the benchmark or launch. THIS WEEK'S 45-MINUTE BLOCK Run one Trajectory Receipt Drill against an agent workflow or evaluation that sits near production. Answer nine questions: 1. What is the exact objective, and which shortcuts remain prohibited even if they improve the score? 2. What configuration differs from normal production use? 3. Where is the hard environment boundary, and what proves isolation? 4. Which identities, credentials, and data sources exist in the run? 5. What action trace is retained across tool calls, permission decisions, retries, environment changes, boundary contacts, and human interventions? 6. Which pattern stops the run? 7. Who has independent stop authority, and has the mechanism been tested? 8. If another person or organization is touched, how are evidence, notice, containment, impact, and remediation handled? 9. What is the final disposition: accepted, rejected, contained, rolled back, remediated, or still under investigation? Score the workflow green, yellow, or red. Do not expand a red workflow. Fix the boundary first, then rerun the drill. LISTENER QUESTION Can your team reconstruct not only what the agent produced, but the full path it took - including the moment someone should have stopped it? SOURCES • OpenAI, OpenAI and Hugging Face partner to address security incident during model evaluation: https://openai.com/index/hugging-face-model-evaluation-security-incident/ • Hugging Face, Security incident disclosure - July 2026: https://huggingface.co/blog/security-incident-july-2026 • OpenAI, Safety and alignment in an era of long-horizon models: https://openai.com/index/safety-alignment-long-horizon-models/ • OpenAI, Introducing OpenAI Presence: https://openai.com/index/introducing-openai-presence/ • OpenAI, Launching Health in ChatGPT: https://openai.com/index/health-in-chatgpt/ • Associated Press incident reporting: https://apnews.com/article/openai-gpt56-sol-hugging-face-63ab84fed5612af04d8a160d60f6def3 LISTEN AND FOLLOW • AI Change De...

AI Change Desk | EP037: Work Agent Receipt Check

A weekly commute-length operator review of the evidence organizations need before delegated AI work can scale.

AI Change Desk | EP036: Preview Before Power Mode

Frontier capability is arriving before broad access. This episode turns OpenAI's GPT-5.6 Sol preview, OpenAI's agent-work research, Microsoft's Copilot in Excel finance controls, and Anthropic's Claude Tag into one operator test: before power mode becomes normal mode, name the preview gate. Who gets frontier capability access, in which surface, with which data and tools, under which safeguards, with what evidence, and who can roll it back? Eligible user group Approved surface Data boundary Tool boundary Safeguard routing Spend and cache budget Evidence log Rollback and deprovision owner OpenAI: https://openai.com/index/previewing-gpt-5-6-sol/ OpenAI Help Center: https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna OpenAI agents/work research: https://openai.com/index/how-agents-are-transforming-work/ Microsoft Copilot in Excel: https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/25/copilot-in-excel-built-for-the-era-of-frontier-finance/ Anthropic Claude Tag: https://www.anthropic.com/news/introducing-claude-tag AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice.

AI Change Desk | EP034: Patch Before Prod

AI security work is moving from "find the bug" toward "help draft the fix." In EP034 of AI Change Desk, Michael breaks down OpenAI's June 22 Daybreak and Patch the Planet announcements and turns them into a practical operator question: If AI can find vulnerabilities and draft fixes at machine speed, who validates the patch, approves the rollout, owns rollback, and proves the fix should ship? Run one 45-minute Patch Before Prod review. Use eight receipts: Finding validator Patch approver Test evidence Release owner Rollback owner Disclosure owner Budget owner Replacement path OpenAI Daybreak: https://openai.com/index/daybreak-securing-the-world/ OpenAI Patch the Planet: https://openai.com/index/patch-the-planet/ OpenAI enterprise spend controls: https://openai.com/index/chatgpt-enterprise-spend-controls/ OpenAI ChatGPT release notes: https://help.openai.com/en/articles/6825453-chatgpt-release-notes Microsoft Copilot Cowork GA: https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/ Microsoft Work IQ APIs: https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/02/announcing-the-new-work-iq-apis/ Anthropic Fable/Mythos access statement: https://www.anthropic.com/news/fable-mythos-access YouTube AI labels update: https://blog.youtube/news-and-events/improving-ai-labels-viewers-creators/ Podnews on RSS AI disclosure flag: https://podnews.net/update/bumper-free Production disclosure: AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval remain human-led. This is operational guidance, not legal advice.

AI Change Desk | EP033: Agent Runtime Budget Check

Bonus
Agent work is not just an access question anymore. It is becoming a runtime budget question. This brief looks at Microsoft's Copilot Cowork general availability and Work IQ controls, then pairs that with Anthropic's statement that it is removing access to Fable 5 / Mythos 5 to ask a practical operator question: If an agent can retrieve context, call tools, run longer tasks, and consume metered credits, who owns the budget gate before the work continues? Run one Agent Runtime Budget Gate by June 24, 2026. Use: Microsoft Copilot Cowork GA: https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/ Microsoft Work IQ APIs: https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/02/announcing-the-new-work-iq-apis/ Anthropic Fable/Mythos access-removal statement: https://www.anthropic.com/news/fable-mythos-access OpenAI ChatGPT release notes: https://help.openai.com/en/articles/6825453-chatgpt-release-notes OpenAI Memory FAQ: https://help.openai.com/en/articles/8590148-memory-faq/ Production disclosure: AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval remain human-led. This is operational guidance, not legal advice.

AI Change Desk | EP031: Memory Control Plane Check

AI memory is becoming more useful, but usefulness creates a new operating surface. If the system can carry context forward, teams need a memory control plane: summary, source, correction, deletion, sensitive-work mode, and disclosure. Why better memory is not just personalization; it is source-of-truth pressure. What OpenAI's June 4 memory rollout changes for operators. Why memory summaries, source tracing, correction, and deletion paths matter. How Lockdown Mode fits sensitive browsing and hostile-input workflows. Why audience disclosure still belongs in the release workflow. A 45-minute memory-control-plane check for Monday teams. What would have to be true for your team to trust remembered AI context in production work? OpenAI memory rollout: https://openai.com/index/chatgpt-memory-dreaming/ OpenAI Memory FAQ: https://help.openai.com/en/articles/8590148-memory-faq/ ChatGPT release notes: https://help.openai.com/en/articles/6825453-chatgpt-release-notes Lockdown Mode: https://help.openai.com/en/articles/20001061-lockdown-mode YouTube AI labels: https://blog.youtube/news-and-events/improving-ai-labels-viewers-creators/ Podnews AI disclosure guide: https://podnews.net/update/ai-disclosures AI-assisted tools were used in parts of the research and production workflow. Final editorial judgment, risk posture, and release approval stayed human-led. This is operational guidance, not legal advice. These are my opinions and are not representative of any organization.
3 de 5