From Muon to Gradient Clipping: Some Thoughts on QK Stability
Summary
A technical exploration of the Muon optimizer's stability when applied to QK updates in Transformer attention. The author traces theoretical issues, attempts principled corrections (including decoupling, pseudoinverses, and a unified bilinear view), explores practical heuristics like gradient clipping, and connects the discussion to low-rank factorization and LoRA.