sergqwer/strata-nvfp4: Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more. Setup installs our GPTQ quants (h
Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more.
ProofCode ↗
Key points
- Setup installs our GPTQ quants (huihui abliterated, OrcaRouter) from Hugging Face.
- CPU experts on AVX-512 or AVX2 (+AVX-VNNI), all experts in RAM from 92 GB, a low-RAM mode below, vision, W4A8 prompts.
- Topics: avx2, avx512, blackwell, cuda, gptq, llm-inference, moe, nvfp4, qwen, vision, windows.
Sources (1)
- [1]sergqwer/strata-nvfp4: Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more. Setup installs our GPTQ quants (hRising AI repositories on GitHub · Sep 29, 04:13 PM
Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more.
Setup installs our GPTQ quants (huihui abliterated, OrcaRouter) from Hugging Face.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 28, 2026Holo4: powering generalist computer-use agents
- Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
- Jul 3, 2026huggingface/transformers v5.13.0: Release v5.13.0
- Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model
