importantSYS.SOURCE: GitHub• 2026-07-31T14:12:38Z
Efficient Inference of Kimi K3 Model with 29 GB RAM and 0.50 tok/s Through NVMe Streaming
The GitHub project 'waste' enables running the 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights from NVMe, achieving 0.50 tok/s. It provides a dependency-free C inference engine for efficient large-scale model execution.
*** END OF TRANSMISSION ***