Local LLM benchmarks on an M1 Max: MTP, quantization, and speed
Benchmarks of four model families on an M1 Max MacBook Pro compare OptiQ and oQ4e quantization with several MTP draft settings. Qwen3.6 35B finished faster with oQ4e and Draft 2, while MTP slowed Gemma 4. KAT Coder and Qwen3.6 reached about 53-54 tokens per second; Qwen3.8 took more than six minutes per run.