Skip to visualization
visualization / Latent · MLA

minGPT · Multi-Head Latent Attention

compressed latent KV state
◆ MLA
T=8
i32
iGated + Query LoRA

🔍 Variable Inspector

click any code line
Click a code line to inspect its variables.
About this lesson
Quick answer

What Multi-Head Latent Attention explains

Cache a compact latent per block and reconstruct keys and values only when attention needs them.

Key concepts

Primary research

DeepSeek-V2 · DeepSeek-AI