AION
Open-source releaseEfficiency & Inference · MLOps, Tooling & Infrastructure1 source · Sep 29, 2026

NVIDIA/TensorRT-LLM v1.3.0rc29

Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151

Key points

  • Support Qwen3.5 checkpoints with global FP8 scales across model loading paths #19519
  • Optimize Gemma4 vision rotary embeddings without a GEMM operation #19571
  • Extend MiniMax-M3 piecewise CUDA graphs through model and executor paths #19423
  • Enable Helix speculative verification for FP8, FP4 MLA, and DSpark groups #19273

Sources (1)

  • [1]NVIDIA/TensorRT-LLM v1.3.0rc29
    GitHub: NVIDIA/TensorRT-LLM · Sep 29, 10:56 AM
    - Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
    - Support Qwen3.5 checkpoints with global FP8 scales across model loading paths #19519

Extractive summary: sentences quoted from the sources.