AION
Opinion / analysisEfficiency & Inference · Hardware & Compute1 source · Oct 10, 2026

Strata with Qwen3.8 Flash Next UD-Q4_K_XL

Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.

Key points

  • Been seeing quite a few posts about Strata lately, so I figured I'd give it a shot on my RTX PRO 6000 and I am very impressed.
  • GPU: 1x RTX PRO 6000 Workstation (96GB)
  • Task Strata (RTX PRO 6000) vLLM (RTX PRO 6000) vLLM (DGX Spark TP2)
  • Qwen 3.8 flash next is flying through coding tasks and it's crazy how efficient it is.

Sources (1)

  • [1]Strata with Qwen3.8 Flash Next UD-Q4_K_XL
    r/LocalLLaMA (top, daily) · Oct 10, 09:50 AM
    Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4_K_XL performs instead.
    Been seeing quite a few posts about Strata lately, so I figured I'd give it a shot on my RTX PRO 6000 and I am very impressed.

Extractive summary: sentences quoted from the sources.