Open sourceOpinion / analysisHardware & Compute · Efficiency & Inference1 source · Sep 29, 2026

sergqwer/strata-nvfp4: Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more. Setup installs our GPTQ quants (h

Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more.

Key points

  • Setup installs our GPTQ quants (huihui abliterated, OrcaRouter) from Hugging Face.
  • CPU experts on AVX-512 or AVX2 (+AVX-VNNI), all experts in RAM from 92 GB, a low-RAM mode below, vision, W4A8 prompts.
  • Topics: avx2, avx512, blackwell, cuda, gptq, llm-inference, moe, nvfp4, qwen, vision, windows.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 28, 2026Holo4: powering generalist computer-use agents
  2. Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
  3. Jul 3, 2026huggingface/transformers v5.13.0: Release v5.13.0
  4. Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0
  5. Jun 10, 2026DiffusionGemma: 4x faster text generation
  6. Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Related