All examples

How DeepSeek shrank the KV cache

The entire input
Multi-Head Latent Attention, explained for engineers
Runtime
1:26

A technical explainer on Multi-Head Latent Attention and why it cuts the KV cache DeepSeek has to hold in memory. Script, narration, animated diagrams and thumbnail all generated from a single line of prompt.

The script, the voiceover, the animation for every scene and the cover thumbnail were all generated. Nothing here was opened in an editor.