< BACK TO NEWS
importantSYS.SOURCE: Hacker News2026-08-04T19:44:55Z

Maple-Preview: Ternary 20B Mixture-of-Experts Model Optimized for iPhone Inference

Maple-Preview is a 20B-parameter ternary Mixture-of-Experts (MoE) model optimized for mobile inference, achieving 120 tokens per second on an iPhone. The project highlights advancements in efficient large model deployment, enabling high-performance AI on consumer hardware.

Comments

Read original article

*** END OF TRANSMISSION ***