48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)
Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
Key points
- It's an 8 bits AWQ + GPTQ quant.
- Made with love and precision (read the model card) on my 4 RTX 3090, I believe it's the best compromise you can actually get on 48Gb of VRAM.
- Running it with vLLM and MTP on 2 RTX 3090 in my workflows and it's working flawlessly at almost 100 tok/s decode (with MTP=5), while being less an overthinker than the OG 3.8 27B :)
- Around 21.6Gb used per card (vLLM launched with 128k context and FP16 KV cache).
Sources (1)
- [1]48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)r/LocalLLaMA (top, daily) · Oct 10, 04:11 PM
Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
It's an 8 bits AWQ + GPTQ quant.
Extractive summary: sentences quoted from the sources.