NVIDIA/TensorRT-LLM v1.3.0rc29
Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
Key points
- Support Qwen3.5 checkpoints with global FP8 scales across model loading paths #19519
- Optimize Gemma4 vision rotary embeddings without a GEMM operation #19571
- Extend MiniMax-M3 piecewise CUDA graphs through model and executor paths #19423
- Enable Helix speculative verification for FP8, FP4 MLA, and DSpark groups #19273
Sources (1)
- [1]NVIDIA/TensorRT-LLM v1.3.0rc29GitHub: NVIDIA/TensorRT-LLM · Sep 29, 10:56 AM
- Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Support Qwen3.5 checkpoints with global FP8 scales across model loading paths #19519
Extractive summary: sentences quoted from the sources.