AnalysisTutorial / explainerEfficiency & Inference · Large Language Models1 source · Oct 11, 2026

OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!

I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!

Proof1 community thread

Key points

  • Yes, there are other 3bit quants for 64GB Mac, but 4bit is the lowest quant I'd tolerate.
  • I manually quantized the original model from the Qwen repo to OQ4E using oMLX, and it turned out to be about 4.72 BPW.
  • It feels like some kind of sorcery to be able to run a 100GB model with 58GB allocated to GPU!
  • Here are the settings I used on oMLX:

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 9, 2026Microsoft's Decision-1 model enters the fast-growing AI decision model race
  2. Oct 9, 2026Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
  3. Oct 8, 2026Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
  4. Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
  5. Oct 8, 2026REMORY: Learning Residual Memory for Context Compaction
  6. Oct 7, 2026Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

Related