AION
Open-source releaseEfficiency & Inference1 source · Oct 8, 2026

huggingface/trl v1.15.0

SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.

Key points

  • ⚡ Fused LM head: up to 6.9× longer sequences on the same GPU
  • Max trainable sequence length (gemma-3-1b, 262k vocabulary, 79 GiB of GPU memory, bf16, batch size 1, gradient checkpointing, sdpa):
  • | Trainer | Peak GiB v1.14.2 → v1.15.0 | Tokens/s |
  • Scoring needs Triton on a GPU (Linux with CUDA, ROCm or XPU).

Sources (1)

  • [1]huggingface/trl v1.15.0
    GitHub: huggingface/trl · Oct 8, 07:26 PM
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a **fused LM head**: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the `[batch, seq, vocab]` logits tensor is never built.
    ### ⚡ Fused LM head: up to 6.9× longer sequences on the same GPU

Extractive summary: sentences quoted from the sources.