The GenAI Evolution Atlas

The GenAI Evolution Atlas

by Peter Liu
Season 2
GenAI News Roundup — week of Sep 5–11
This week: OpenAI's contested claim that an agent swarm cracked the Navier–Stokes Millennium Prize Problem draws a same-day credit dispute, even as Anthropic's Claude agents post a genuinely verified feat — a machine-checked proof of Fermat's Last Theorem; Google, Anthropic, and OpenAI simultaneously roll out cyber-focused models and safeguard programs; and DeepSeek's smaller V4.1-Flash gets its own larger V4-Pro pulled from production — plus commentary and what's next. Covers: the Fermat's Last Theorem formalization; Terminal-Universe; LLaDA-Image; on-policy distillation with one training example; OpenAI's automated research intern milestone; a critical review of agentic AI progress; Compile by Training; An Alien Mind; Don't Drop Dropout; CUA-Universe; the Navier–Stokes claim and dispute; fractal basins in latent reasoning; contextual understanding evaluation; MoE expert-halving; self-consensus early-exit risks; structural process supervision for latent CoT; distribution-consistent MoE inference; VLX-VR; DeepSeek V4.1-Flash; the Google/Anthropic/OpenAI cyber safeguards announcement; OpenAI's Agents API and finance workspace; and Sebastian Raschka's looped-transformer explainer.
Deep Dive: Claude Agents Formalize Fermat's Last Theorem
Working largely autonomously over 11 days on the open Prove2Me platform, many coordinated Claude agents produced the first complete, machine-checked proof of Fermat's Last Theorem in Lean 4 — over 13 million lines of code and roughly 29,500 new theorems, dwarfing Lean's existing main math library and closing out the 20-year-old Wiedijk "100 theorems" formalization challenge list. This episode digs into how a proof of that scale gets built and verified by a swarm of AI agents, why a machine-checked result is such a hard-to-fake data point, and what it does and doesn't tell us about the state of long-horizon autonomous AI work. Source: https://www.anthropic.com/research/formalizing-fermats-last-theorem
GenAI News Roundup — week of Aug 31–Sep 4
This week: OpenAI's Astra crosses the "Critical" cyber capability threshold and then actually ships as GPT-6 Astra, NVIDIA moves to acquire Hugging Face for ~$13B days after OpenAI's own Hugging Face security-incident report, and labs keep converging on the same sparse-MoE + hybrid-attention playbook — plus commentary and what's next. Covers: the OpenAI Hugging Face incident report; GLM-5.3-Flash and its architectural convergence with Qwen3.8-Flash-Next; Anthropic's automated-alignment-researcher results; DeepSeek V4-Pro's GA; ContextPilot and PLVR; OpenAI's Astra Critical-threshold announcement and the GPT-6 Astra launch; Claude Fable 5.1 and Mythos 5.1; Gemini 3.8 Flash; Qwen3.8-Max-0902; Latent Recurrent Thoughts; MASkills; thinking-effort alignment in abductive reasoning; NVIDIA's Hugging Face acquisition; NVIDIA's gold-medal competitive-programming post-training work; and a statistical theory of Mixture-of-Experts.
Deep Dive: When AI Crosses the Critical Line — Inside OpenAI's Astra
AI
OpenAI's Astra is the first model ever assessed as reaching "Critical" cyber capability under the company's Preparedness Framework — able to find and exploit previously-unknown zero-days in hardened systems and plan/execute end-to-end cyberattacks from only a high-level goal, without step-by-step human guidance. This episode digs into what that threshold means, how a capability like this gets evaluated and gated, and why Astra still ships — just with tightened access controls and monitoring rather than staying on the shelf. Source: https://openai.com/index/path-to-astra/
Season 1
S1E8: Why AI Now Reasons and Acts
AI
Frontier systems & the engineered stack From "a model" to "a system" — retrieval, tools, agents, and reasoning.
S1E7: How AI Gained Eyes and Ears
AI
Multimodality & generation beyond text Give models eyes and ears — and learn to generate pixels.
S1E6: How alignment turned autocomplete into AI assistants
AI
Alignment & post-training Turn a next-token predictor into a helpful, honest assistant.
S1E5: How FlashAttention and MoE saved AI scaling
AI
Efficiency & better building blocks Make big Transformers faster, longer, and cheaper to run.
S1E4: How massive scale triggered emergent AI
AI
Scale and emergence Make the same architecture enormous — and new behaviour appears.
S1E3: The Big Bang of Modern AI
AI
The pretraining era Pretrain once on a mountain of text, then transfer everywhere.
1 of 2