huggingface/trl v1.14.2
Patch release fixing two cases of silently wrong training and three crashes.
Key points
- Worth calling out: #7505 affects GRPO, RLOO and Distillation on the default generation path (usevllm=False).
- Models that end a turn with an id only their generation config declares (Gemma 3/4, Phi-3.5, ...) did not stop at the end of the turn, and everything generated after it was trained on, up to maxcompletionlength.
- With vLLM, generation stopped correctly but the metrics were still wrong, and masktruncatedcompletions=True dropped every finished completion.
- Runs on models whose tokenizer eos is their only end-of-turn id are unchanged.
Sources (1)
- [1]huggingface/trl v1.14.2GitHub: huggingface/trl · Oct 6, 08:30 PM
Patch release fixing two cases of silently wrong training and three crashes.
Worth calling out: **#7505** affects GRPO, RLOO and Distillation on the default generation path (`use_vllm=False`).
Extractive summary: sentences quoted from the sources.