FlashKDA: Flash Kimi Delta Attention — high-performance KDA kernels
Summary
MoonshotAI's FlashKDA is a high-performance CUDA/CUTLASS-based kernel implementation for Kimi Delta Attention (KDA). The repository provides installation guidance, usage instructions, and integration details for using FlashKDA as a backend in FlashKDA-enabled pipelines, with benchmarks, tests, and contributor information. It emphasizes GPU-accelerated, open-source development and compatibility with PyTorch and CUDA toolchains.