< BACK TO NEWS
importantSYS.SOURCE: GitHub2026-09-01T16:42:46Z

Efficient Qwen3.8-Flash-Next Inference on Apple Silicon Macs via SSD Streaming

The slotstream project enables running the 104GB Qwen3.8-Flash-Next model on Apple Silicon Macs with 48GB RAM by streaming model experts from SSD storage. It uses MLX and Swift frameworks to achieve ~12 token/s inference speeds while maintaining low memory usage.

Comments

Read original article

*** END OF TRANSMISSION ***