importantSYS.SOURCE: Hacker News• 2026-08-04T19:44:55Z
Maple-Preview: Ternary 20B Mixture-of-Experts Model Optimized for iPhone Inference
Maple-Preview is a 20B-parameter ternary Mixture-of-Experts (MoE) model optimized for mobile inference, achieving 120 tokens per second on an iPhone. The project highlights advancements in efficient large model deployment, enabling high-performance AI on consumer hardware.
*** END OF TRANSMISSION ***