Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
Key points
- Results (same hardware, my megakernel vs llama.cpp with MTP):
- Writing code: 140 tok/s vs 73 (1.9x faster)
- The speedup comes from running each whole speculative-decoding cycle in one kernel launch, so checking 4-5 drafted tokens costs about the same as checking one.
- It's an OpenAI-compatible server, so it works as a drop-in replacement for the llama server.
Sources (1)
- [1]Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernelr/LocalLLaMA (top, daily) · Oct 10, 01:01 PM
I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
Results (same hardware, my megakernel vs llama.cpp with MTP):
Extractive summary: sentences quoted from the sources.