AION
Open-source releaseReinforcement Learning · Large Language Models · Safety & Alignment1 source · Oct 6, 2026

huggingface/trl v1.14.2

Patch release fixing two cases of silently wrong training and three crashes.

Key points

  • Worth calling out: #7505 affects GRPO, RLOO and Distillation on the default generation path (usevllm=False).
  • Models that end a turn with an id only their generation config declares (Gemma 3/4, Phi-3.5, ...) did not stop at the end of the turn, and everything generated after it was trained on, up to maxcompletionlength.
  • With vLLM, generation stopped correctly but the metrics were still wrong, and masktruncatedcompletions=True dropped every finished completion.
  • Runs on models whose tokenizer eos is their only end-of-turn id are unchanged.

Sources (1)

  • [1]huggingface/trl v1.14.2
    GitHub: huggingface/trl · Oct 6, 08:30 PM
    Patch release fixing two cases of silently wrong training and three crashes.
    Worth calling out: **#7505** affects GRPO, RLOO and Distillation on the default generation path (`use_vllm=False`).

Extractive summary: sentences quoted from the sources.