S1E5: How FlashAttention and MoE ...
AI
S1E5: How FlashAttention and MoE saved AI scaling
AI

The GenAI Evolution Atlas by Peter Liu

Episode notes

Efficiency & better building blocks

Make big Transformers faster, longer, and cheaper to run.

Keywords
Generative AILLMAttentionKV cacheMoEFlashAttention