AION
Opinion / analysisEfficiency & Inference1 source · Oct 10, 2026

Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel

I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.

Key points

  • Results (same hardware, my megakernel vs llama.cpp with MTP):
  • Writing code: 140 tok/s vs 73 (1.9x faster)
  • The speedup comes from running each whole speculative-decoding cycle in one kernel launch, so checking 4-5 drafted tokens costs about the same as checking one.
  • It's an OpenAI-compatible server, so it works as a drop-in replacement for the llama server.

Sources (1)

  • [1]Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
    r/LocalLLaMA (top, daily) · Oct 10, 01:01 PM
    I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
    Results (same hardware, my megakernel vs llama.cpp with MTP):

Extractive summary: sentences quoted from the sources.