What Actually Predicts Local LLM Speed on a 32 GB Mac (2026): I Benchmarked 7 Models
The short version: On a Mac mini (Apple M2 Pro, 32 GB), the fastest model I tested was also the biggest — Qwen3-Coder 30B, a Mixture-of-Experts model with only ~3B active parameters, ran at 42.8 tokens/sec, beating a plain 8B model. Total parameter count barely predicts speed on this hardware; active parameters do. And the much-hyped … Read more