Open source
Libraries, tools and repos: releases you depend on and projects gaining stars.
huggingface/trl v1.15.0
SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
Converting dense models into Mixture-of-Experts
For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
BerriAI/litellm v1.105.0
Verify using the pinned commit hash (recommended):
Faster Scientific Image Analysis with NVIDIA cuPhoton
Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can...Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can process it to support timely decisions.
unslothai/unsloth v0.1.905-beta: Sandboxing is here!
We're introducing Windows, Mac and Linux sandboxing in Unsloth!
vllm-project/vllm v0.31.0
Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit
Mia-AiLab published the model GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit on Hugging Face.
pydantic/pydantic-ai clai2-bleeding: CLAI2 bleeding
Sdists rebuilt from main for /update on the CLAI bleeding channel.
NVIDIA/TensorRT-LLM v1.3.0rc29
Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
pytorch/pytorch v2.14.1: PyTorch 2.14.1 Release
This release is meant to fix the following regressions and silent correctness issues:
sergqwer/strata-nvfp4: Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more. Setup installs our GPTQ quants (h
Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more.
ollama/ollama v0.34.4
Qwen 3.8 prompt processing is faster on Apple Silicon.
Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw
Infatoshi published the model GLM-5.3-UNCENSORED-EXL3-3.0bpw on Hugging Face.
BerriAI/litellm v1.104.3
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.104.2
Verify using the pinned commit hash (recommended):
AI Native by Design: Lessons Learned from Building NVIDIA TensorRT Model Connect
Parallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT...Parallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT Model Connect is an open source collection of AI model reference implementations in C++, built on top of NVIDIA TensorRT.
BerriAI/litellm v1.102.4
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.101.6
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.104.1
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.103.4
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.102.3
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.100.5
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.101.5
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.104.0
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.103.3
Verify using the pinned commit hash (recommended):
unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
Run and serve Decision Models like Laya (open-source Jev) locally
BerriAI/litellm v1.103.2
Verify using the pinned commit hash (recommended):
BerriAI/litellm v1.101.4
Verify using the pinned commit hash (recommended):
vllm-project/vllm v0.30.0
This release features 762 commits from 315 contributors (104 new)!
Dreamer-Toby/STEPQuant: STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization Language: Python.
StayLameBro/backburner: Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable
Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable Language: Python.