Attention Through Arithmetic Intensity
Summary
The article analyzes attention architectures (MHA, GQA, MQA, MLA) through arithmetic intensity, deriving how AI scales with head counts and caching strategies. It introduces MLA's latent cache and contrasts decode vs prefill, explores implications for MTP and sparsity, and provides practical guidance for inference optimization and deployment.