importantSYS.SOURCE: GitHub• 2026-09-01T16:42:46Z
Efficient Qwen3.8-Flash-Next Inference on Apple Silicon Macs via SSD Streaming
The slotstream project enables running the 104GB Qwen3.8-Flash-Next model on Apple Silicon Macs with 48GB RAM by streaming model experts from SSD storage. It uses MLX and Swift frameworks to achieve ~12 token/s inference speeds while maintaining low memory usage.
*** END OF TRANSMISSION ***