huggingface/trl v1.15.0
SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
Key points
- ⚡ Fused LM head: up to 6.9× longer sequences on the same GPU
- Max trainable sequence length (gemma-3-1b, 262k vocabulary, 79 GiB of GPU memory, bf16, batch size 1, gradient checkpointing, sdpa):
- | Trainer | Peak GiB v1.14.2 → v1.15.0 | Tokens/s |
- Scoring needs Triton on a GPU (Linux with CUDA, ROCm or XPU).
Sources (1)
- [1]huggingface/trl v1.15.0GitHub: huggingface/trl · Oct 8, 07:26 PM
SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a **fused LM head**: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the `[batch, seq, vocab]` logits tensor is never built.
### ⚡ Fused LM head: up to 6.9× longer sequences on the same GPU
Extractive summary: sentences quoted from the sources.