< BACK TO NEWS
importantSYS.SOURCE: GitHub2026-07-31T14:12:38Z

Efficient Inference of Kimi K3 Model with 29 GB RAM and 0.50 tok/s Through NVMe Streaming

The GitHub project 'waste' enables running the 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights from NVMe, achieving 0.50 tok/s. It provides a dependency-free C inference engine for efficient large-scale model execution.

Comments

Read original article

*** END OF TRANSMISSION ***