Alex Leko
All content on this blog was fully or partially created using local AI (Apple MLX).

Local LLM benchmarks on an M1 Max: MTP, quantization, and speed

  • llm
  • benchmark
  • apple
  • quantization
  • mtp
  • m1max
  • machinelearning
  • localai
  • performance
  • trending
  • featured

I benchmarked four model families on an M1 Max using OptiQ and oQ4e quantization, with several Multi-Token Prediction (MTP) draft settings.

The results varied by model. Qwen3.6 35B with oQ4e and MTP Draft 2 finished in 56.7 seconds, down from 95.5 seconds without MTP. Gemma 4 26B slowed down with MTP: its oQ4e Draft 4 run took 114.0 seconds, compared with 51.3 seconds without MTP. Draft 4 added no clear benefit over Draft 2 in most of the other tests. KAT Coder V2.5 and Qwen3.6 reached about 53-54 tokens per second. Qwen3.8 took more than six minutes per run, mostly in its reasoning phase.

Test Setup & Methodology

I ran all benchmarks locally with oMLX on unified memory:

  • Hardware: Apple MacBook Pro (M1 Max, 64GB Unified Memory)
  • Prompt: "Explain to me the ports and adapters architecture pattern in Python"
  • Evaluation Metrics: Prefill speed (t/s), Token Generation speed (t/s), Thinking time (seconds), and Total Duration (seconds).

Global Inference Parameters

Model Family Max Tokens Context Window Temp Top P Top K Min P Rep. P Pres. P Thinking Limit
Qwen3.6 8k 75k 0.6 0.95 20 0.0 1.0 0.0 4096
Qwen3.8 8k 36k 1.0 0.95 20 0.0 1.0 0.0 4096
KAT Coder v2.5 8k 75k 0.6 0.95 20 0.0 1.0 0.0 4096
Gemma 4 26B A4B 8k 75k 1.0 0.95 64 — — — 4096

Benchmark Results by Model Family

1. KAT Coder V2.5 Dev

KAT Coder's generation speed stayed between 50.1 and 53.1 tokens per second across these runs.

Quantization MTP Draft Prefill (t/s) Gen Speed (t/s) Thinking (s) Total Duration (s)
OptiQ 4-bit None 31.8 50.1 0.6 63.1
oQ4e None 47.1 53.1 0.6 65.7
oQ4e Draft 2 32.1 52.7 0.6 59.5
oQ4e Draft 4 ⭐ 33.0 53.1 0.6 59.0

KAT Coder generated at about 50-53 tokens per second. The oQ4e MTP runs finished in 59.0-59.5 seconds, compared with 65.7 seconds without MTP.

2. Qwen3.6 35B A3B

Qwen3.6's fastest run used oQ4e with MTP Draft 2.

Quantization MTP Draft Prefill (t/s) Gen Speed (t/s) Thinking (s) Total Duration (s)
OptiQ 4-bit None 29.1 45.2 44.1 79.7
OptiQ 4-bit Draft 2 53.4 51.9 36.1 80.3
OptiQ 4-bit Draft 4 32.3 52.6 35.6 73.5
oQ4e None 35.1 37.7 47.5 95.5
oQ4e Draft 2 ⭐ 35.1 54.1 30.0 56.7

With oQ4e, Draft 2 raised generation speed from 37.7 to 54.1 tokens per second and reduced total duration from 95.5 to 56.7 seconds.

3. Gemma 4 26B A4B

Gemma 4 slowed down with each tested MTP draft setting.

Quantization MTP Draft Prefill (t/s) Gen Speed (t/s) Thinking (s) Total Duration (s)
OptiQ 4-bit None 18.9 40.6 20.8 58.7
OptiQ 4-bit Draft 2 23.6 36.6 23.0 67.9
OptiQ 4-bit Draft 4 29.3 31.8 27.5 72.8
oQ4e None ⭐ 17.2 41.7 14.6 51.3
oQ4e Draft 2 34.4 35.8 23.1 65.1
oQ4e Draft 4 24.6 20.5 39.1 114.0

Gemma 4's fastest run used oQ4e without MTP and took 51.3 seconds. With Draft 4, generation speed fell to 20.5 tokens per second and total duration rose to 114.0 seconds.

4. Qwen3.8 27B

Qwen3.8 took more than six minutes in each run under this oMLX configuration.

Quantization MTP Draft Prefill (t/s) Gen Speed (t/s) Thinking (s) Total Duration (s)
OptiQ 4-bit None 36.6 10.0 9.8 411.4 (6.85 min)
OptiQ 4-bit Draft 2 37.9 11.2 9.9 371.5 (6.18 min)
OptiQ 4-bit Draft 4 36.7 10.2 9.7 404.7 (6.73 min)
oQ4e None 37.6 11.2 365.4 367.2 (6.12 min)
oQ4e Draft 2 ⭐ 37.3 11.3 362.7 364.5 (6.07 min)
oQ4e Draft 4 36.0 11.1 368.9 370.07 (6.16 min)

Qwen3.8 generated at about 10-11 tokens per second. In the oQ4e runs, roughly 363-369 seconds were recorded as thinking time, close to the configured 4096-token thinking limit. In oMLX chat, the response appeared only after completion. These measurements do not establish why the model was slow; a reasoning loop or missing MLX kernel support may be involved.

Analysis

MTP Mechanics: Acceleration vs. Overhead

MTP generates draft tokens that the main model then verifies. Accepted drafts can reduce the work needed to produce a response; rejected drafts add verification overhead. The benchmark did not measure acceptance rates, so the results show the performance difference but do not explain its cause. Qwen3.6 reached 54.1 tokens per second with Draft 2, while Gemma 4 fell to 20.5 tokens per second with Draft 4.

Quantization Comparison: oQ4e vs. OptiQ

In these tests, oQ4e with a suitable MTP setting produced the fastest runs for Qwen3.6 and KAT Coder. OptiQ speeds varied less across MTP settings for some models, but Qwen3.8 remained slow with either quantization.

Configurations to try

For developers running local inference on an Apple Silicon M1 Max (64GB RAM):

Model Recommended Config Performance Primary Use Case
KAT Coder V2.5 oQ4e + MTP Draft 4 53.1 t/s (59.0s) Fast, reliable code generation
Qwen3.6 35B oQ4e + MTP Draft 2 54.1 t/s (56.7s) Deep reasoning & architectural design
Gemma 4 26B oQ4e (No MTP) 41.7 t/s (51.3s) General purpose & fast response
Qwen3.8 27B Avoid currently 11.3 t/s (364s) Await MLX kernel updates

MTP results varied substantially by model in this test. Benchmark each model and draft setting on your own workload before choosing a default.