AION
Opinion / analysisEfficiency & Inference1 source · Oct 10, 2026

48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)

Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.

Key points

  • It's an 8 bits AWQ + GPTQ quant.
  • Made with love and precision (read the model card) on my 4 RTX 3090, I believe it's the best compromise you can actually get on 48Gb of VRAM.
  • Running it with vLLM and MTP on 2 RTX 3090 in my workflows and it's working flawlessly at almost 100 tok/s decode (with MTP=5), while being less an overthinker than the OG 3.8 27B :)
  • Around 21.6Gb used per card (vLLM launched with 128k context and FP16 KV cache).

Sources (1)

Extractive summary: sentences quoted from the sources.