Local LLM benchmarks on an M1 Max: MTP, quantization, and speed
I benchmarked four model families on an M1 Max using OptiQ and oQ4e quantization, with several Multi-Token Prediction (MTP) draft settings.
The results varied by model. Qwen3.6 35B with oQ4e and MTP Draft 2 finished in 56.7 seconds, down from 95.5 seconds without MTP. Gemma 4 26B slowed down with MTP: its oQ4e Draft 4 run took 114.0 seconds, compared with 51.3 seconds without MTP. Draft 4 added no clear benefit over Draft 2 in most of the other tests. KAT Coder V2.5 and Qwen3.6 reached about 53-54 tokens per second. Qwen3.8 took more than six minutes per run, mostly in its reasoning phase.
Test Setup & Methodology
I ran all benchmarks locally with oMLX on unified memory:
- Hardware: Apple MacBook Pro (M1 Max, 64GB Unified Memory)
- Prompt: "Explain to me the ports and adapters architecture pattern in Python"
- Evaluation Metrics: Prefill speed (t/s), Token Generation speed (t/s), Thinking time (seconds), and Total Duration (seconds).
Global Inference Parameters
| Model Family | Max Tokens | Context Window | Temp | Top P | Top K | Min P | Rep. P | Pres. P | Thinking Limit |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.6 | 8k | 75k | 0.6 | 0.95 | 20 | 0.0 | 1.0 | 0.0 | 4096 |
| Qwen3.8 | 8k | 36k | 1.0 | 0.95 | 20 | 0.0 | 1.0 | 0.0 | 4096 |
| KAT Coder v2.5 | 8k | 75k | 0.6 | 0.95 | 20 | 0.0 | 1.0 | 0.0 | 4096 |
| Gemma 4 26B A4B | 8k | 75k | 1.0 | 0.95 | 64 | — | — | — | 4096 |
Benchmark Results by Model Family
1. KAT Coder V2.5 Dev
KAT Coder's generation speed stayed between 50.1 and 53.1 tokens per second across these runs.
| Quantization | MTP Draft | Prefill (t/s) | Gen Speed (t/s) | Thinking (s) | Total Duration (s) |
|---|---|---|---|---|---|
| OptiQ 4-bit | None | 31.8 | 50.1 | 0.6 | 63.1 |
| oQ4e | None | 47.1 | 53.1 | 0.6 | 65.7 |
| oQ4e | Draft 2 | 32.1 | 52.7 | 0.6 | 59.5 |
| oQ4e | Draft 4 ⭐ | 33.0 | 53.1 | 0.6 | 59.0 |
KAT Coder generated at about 50-53 tokens per second. The oQ4e MTP runs finished in 59.0-59.5 seconds, compared with 65.7 seconds without MTP.
2. Qwen3.6 35B A3B
Qwen3.6's fastest run used oQ4e with MTP Draft 2.
| Quantization | MTP Draft | Prefill (t/s) | Gen Speed (t/s) | Thinking (s) | Total Duration (s) |
|---|---|---|---|---|---|
| OptiQ 4-bit | None | 29.1 | 45.2 | 44.1 | 79.7 |
| OptiQ 4-bit | Draft 2 | 53.4 | 51.9 | 36.1 | 80.3 |
| OptiQ 4-bit | Draft 4 | 32.3 | 52.6 | 35.6 | 73.5 |
| oQ4e | None | 35.1 | 37.7 | 47.5 | 95.5 |
| oQ4e | Draft 2 ⭐ | 35.1 | 54.1 | 30.0 | 56.7 |
With oQ4e, Draft 2 raised generation speed from 37.7 to 54.1 tokens per second and reduced total duration from 95.5 to 56.7 seconds.
3. Gemma 4 26B A4B
Gemma 4 slowed down with each tested MTP draft setting.
| Quantization | MTP Draft | Prefill (t/s) | Gen Speed (t/s) | Thinking (s) | Total Duration (s) |
|---|---|---|---|---|---|
| OptiQ 4-bit | None | 18.9 | 40.6 | 20.8 | 58.7 |
| OptiQ 4-bit | Draft 2 | 23.6 | 36.6 | 23.0 | 67.9 |
| OptiQ 4-bit | Draft 4 | 29.3 | 31.8 | 27.5 | 72.8 |
| oQ4e | None ⭐ | 17.2 | 41.7 | 14.6 | 51.3 |
| oQ4e | Draft 2 | 34.4 | 35.8 | 23.1 | 65.1 |
| oQ4e | Draft 4 | 24.6 | 20.5 | 39.1 | 114.0 |
Gemma 4's fastest run used oQ4e without MTP and took 51.3 seconds. With Draft 4, generation speed fell to 20.5 tokens per second and total duration rose to 114.0 seconds.
4. Qwen3.8 27B
Qwen3.8 took more than six minutes in each run under this oMLX configuration.
| Quantization | MTP Draft | Prefill (t/s) | Gen Speed (t/s) | Thinking (s) | Total Duration (s) |
|---|---|---|---|---|---|
| OptiQ 4-bit | None | 36.6 | 10.0 | 9.8 | 411.4 (6.85 min) |
| OptiQ 4-bit | Draft 2 | 37.9 | 11.2 | 9.9 | 371.5 (6.18 min) |
| OptiQ 4-bit | Draft 4 | 36.7 | 10.2 | 9.7 | 404.7 (6.73 min) |
| oQ4e | None | 37.6 | 11.2 | 365.4 | 367.2 (6.12 min) |
| oQ4e | Draft 2 ⭐ | 37.3 | 11.3 | 362.7 | 364.5 (6.07 min) |
| oQ4e | Draft 4 | 36.0 | 11.1 | 368.9 | 370.07 (6.16 min) |
Qwen3.8 generated at about 10-11 tokens per second. In the oQ4e runs, roughly 363-369 seconds were recorded as thinking time, close to the configured 4096-token thinking limit. In oMLX chat, the response appeared only after completion. These measurements do not establish why the model was slow; a reasoning loop or missing MLX kernel support may be involved.
Analysis
MTP Mechanics: Acceleration vs. Overhead
MTP generates draft tokens that the main model then verifies. Accepted drafts can reduce the work needed to produce a response; rejected drafts add verification overhead. The benchmark did not measure acceptance rates, so the results show the performance difference but do not explain its cause. Qwen3.6 reached 54.1 tokens per second with Draft 2, while Gemma 4 fell to 20.5 tokens per second with Draft 4.
Quantization Comparison: oQ4e vs. OptiQ
In these tests, oQ4e with a suitable MTP setting produced the fastest runs for Qwen3.6 and KAT Coder. OptiQ speeds varied less across MTP settings for some models, but Qwen3.8 remained slow with either quantization.
Configurations to try
For developers running local inference on an Apple Silicon M1 Max (64GB RAM):
| Model | Recommended Config | Performance | Primary Use Case |
|---|---|---|---|
| KAT Coder V2.5 | oQ4e + MTP Draft 4 | 53.1 t/s (59.0s) | Fast, reliable code generation |
| Qwen3.6 35B | oQ4e + MTP Draft 2 | 54.1 t/s (56.7s) | Deep reasoning & architectural design |
| Gemma 4 26B | oQ4e (No MTP) | 41.7 t/s (51.3s) | General purpose & fast response |
| Qwen3.8 27B | Avoid currently | 11.3 t/s (364s) | Await MLX kernel updates |
MTP results varied substantially by model in this test. Benchmark each model and draft setting on your own workload before choosing a default.