Swapping a Dense 27B for a MoE on Apple Silicon
The wait
I run my own inference endpoint on a Mac Studio (M4 Max, 128 GB) — oMLX behind an API-key gateway, serving chat, vision, speech and now images to a handful of my apps. The default chat model was a dense Qwen3.8 27B. Good answers. Painful waits.
Hand it an 8,000-token prompt — a system prompt, a couple of files, some history — and it spent 31.7 seconds reading before it wrote a single word. Then it wrote at 29 tokens a second.

A day of tuning the wrong thing
oMLX has a lot of knobs, and I tried the ones aimed at exactly this:
- Neural Engine prefill — hand prompt processing to the ANE. 33.6 s → 32.5 s. Noise.
- KV cache quantization — 33.3 s → 31.5 s. About 5%, and lossy.
- Speculative prefill — built for mixture-of-experts models. Mine wasn’t one.
I even wrote the conclusion down: ~33 seconds is just what this hardware does.
It wasn’t. It’s what a dense 27B does.
Active parameters are the lever
A dense model runs every parameter for every token. A mixture-of-experts (MoE) model sends each token through a few “experts” and skips the rest. Qwen3.6-35B-A3B holds 35 billion parameters but only uses about 3 billion per token — and that holds for reading your prompt, not just writing the reply.
Same harness, same machine:
| Dense 27B | MoE 35B-A3B | |
|---|---|---|
| Read an 8k-token prompt | 31.7 s | 4.4 s |
| Same prompt, cached | 13.9 s | 2.1 s |
| Generation speed | 29 tok/s | 150 tok/s |
Five to seven times faster everywhere. Both hold a 256K-token context, but only one of them makes it practical to fill.

Faster isn’t the question. Worse is.
A model that’s fast and dumb is just a quicker way to be wrong. So before switching anything, both models answered the same 707 questions, and nothing was graded by vibes or by another model:
- Code (HumanEval+, 163 problems) — the generated code was run against the tests, in a container with no network
- Knowledge (MMLU-Pro, 294 questions) — exact match on the chosen answer
- Math (GSM8K, 250 problems) — exact match on the number
| MoE | Dense 27B | |
|---|---|---|
| Code | 90.8% | 90.8% |
| Knowledge | 67.3% | 64.6% |
| Math | 96.0% | 96.0% |
| Overall | 82.9% | 81.8% |
A tie. The knowledge edge isn’t statistically real (p = 0.27), so I’m not claiming the MoE is better — just that it isn’t worse.
Two things I’d recommend:
- Write the decision rule before you look. Mine was “within 2 points overall, and not significantly worse anywhere.” Otherwise you’ll talk yourself into whatever the numbers say.
- Test the grader. Mine passed all 163 correct reference solutions and failed 163 deliberately wrong ones. A grader that passes everything looks exactly like two perfect models.
Bonus: speculative decoding can make things slower
Both models support multi-token prediction (MTP): the model drafts a few tokens ahead and checks them in one pass. On the MoE that was a 41% speedup. On the dense 27B it made generation 26% slower — even though it accepted more of its drafts.
The difference is the checking pass: about 18 ms on the MoE, about 110 ms on the dense model. Verifying four tokens at once is only cheap while it stays memory-bound, and on 27 billion active parameters it isn’t. Acceptance rate alone tells you nothing — the cost of verifying decides it.
What I’d tell past me
- Change the model’s shape before touching a setting. A day of knobs got me 5%. One swap got me 7×.
- Measure quality mechanically. Run the code, match the answers, pair the questions.
- Only try MTP on MoE models, and time it instead of trusting the acceptance rate.
The dense 27B is retired, my default chat model is the MoE, and every conversation on my stack got several times faster overnight 🎉
