AION
Modelofficial release

nvidia/Qwen3.8-2.4T-A95B-NVFP4

Parameters1261B (95B active)
Context256K
Input → outputtext → text
Licenseother
PrecisionU8 · nvfp4
ArchitectureQwen3_5MoeForCausalLM
ReleasedAug 25, 2026
Downloads10.6K

Based on Qwen/Qwen3.8-2.4T-A95B

Other sizes and formats in the family

ModelParamsPrecisionDownloads
nvidia/Qwen3.8-27B-NVFP4nvfp418.2BU8621.6K

Compare side by side →

This release in the news

No coverage of this exact release yet.

Recent news about Qwen

  1. Oct 11, 2026 · Opinion / analysis
    UPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp
  2. Oct 11, 2026 · Opinion / analysis
    [P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]
  3. Oct 11, 2026 · Opinion / analysis
    PSA: DeepSeek V4.1 Flash habitually exfiltrates API keys. It is dangerously misaligned and may be hazardous to use
  4. Oct 11, 2026 · Opinion / analysis
    Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis
  5. Oct 11, 2026 · Open-source release
    Converting dense models into Mixture-of-Experts
  6. Oct 11, 2026 · Tutorial / explainer
    Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards
  7. Oct 11, 2026 · Opinion / analysis
    Qwen3.8 Flash Next fixed my GNOME extension
  8. Oct 11, 2026 · Tutorial / explainer
    Building a 4x R9700 setup for a 10 person startup