DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Kimi Linear: An Expressive, Efficient Attention Architecture

Quality: 9/10 Relevance: 9/10

Summary

Kimi Linear: An Expressive, Efficient Attention Architecture introduces Kimi Linear, a hybrid linear attention architecture that outperforms full attention in various contexts. The paper presents Kimi Delta Attention (KDA), a chunkwise algorithm with a Diagonal-Plus-Low-Rank variant, and reports substantial reductions in KV cache usage and improvements in decoding throughput, while also open-sourcing the KDA kernel and vLLM implementations and releasing pretrained model checkpoints.

🚀 Service construit par Johan Denoyer