OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!
I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
Proof1 community thread
Key points
- Yes, there are other 3bit quants for 64GB Mac, but 4bit is the lowest quant I'd tolerate.
- I manually quantized the original model from the Qwen repo to OQ4E using oMLX, and it turned out to be about 4.72 BPW.
- It feels like some kind of sorcery to be able to run a 100GB model with 58GB allocated to GPU!
- Here are the settings I used on oMLX:
Sources (1)
- [1]OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!r/LocalLLaMA (top, daily) · Oct 11, 12:00 AM
I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
Yes, there are other 3bit quants for 64GB Mac, but 4bit is the lowest quant I'd tolerate.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 9, 2026Microsoft's Decision-1 model enters the fast-growing AI decision model race
- Oct 9, 2026Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
- Oct 8, 2026Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
- Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
- Oct 8, 2026REMORY: Learning Residual Memory for Context Compaction
- Oct 7, 2026Cache the Encoder Within:Compact, Reusable Memory across LLM Queries