Big vs. Small Local Models on a 32 GB Mac: Where MoE Breaks the ‘Bigger = Slower’ Rule
BLUF On my 32 GB Mac mini (M2 Pro), going from an 8B dense model to a 14B dense model drops generation throughput from 26.1 tok/s to 14.7 tok/s — a real, felt slowdown. But two MoE models, qwen3-coder:30b and gpt-oss:20b, hit 40.0 tok/s and 32.8 tok/s respectively, both faster than the little 8B Llama … Read more