positiveSYS.SOURCE: arXiv.org• 2026-07-28T10:52:30Z
Kimi Linear: A Hybrid Linear Attention Architecture for Enhanced Efficiency and Performance
Kimi Linear introduces a hybrid linear attention architecture that outperforms full attention across multiple scenarios while reducing KV cache usage by 75% and achieving 6x faster decoding throughput. It employs Kimi Delta Attention (KDA) with a refined gating mechanism and optimized DPLR matrices for hardware efficiency.
*** END OF TRANSMISSION ***