Open source

Libraries, tools and repos: releases you depend on and projects gaining stars.

GitHub: huggingface/trl3d ago

huggingface/trl v1.15.0

SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.

Code
r/LocalLLaMA (top, daily)17h ago

Converting dense models into Mixture-of-Experts

For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.

GitHub: BerriAI/litellm13h ago

BerriAI/litellm v1.105.0

Verify using the pinned commit hash (recommended):

Code
NVIDIA Technical Blog4d ago

Faster Scientific Image Analysis with NVIDIA cuPhoton

Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can...Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can process it to support timely decisions.

GitHub: unslothai/unsloth3d ago

unslothai/unsloth v0.1.905-beta: Sandboxing is here!

We're introducing Windows, Mac and Linux sandboxing in Unsloth!

Code
GitHub: vllm-project/vllm6d ago

vllm-project/vllm v0.31.0

Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).

Code
Hugging Face trending models6d ago

Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit

Mia-AiLab published the model GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit on Hugging Face.

Weights
GitHub: pydantic/pydantic-ai5d ago

pydantic/pydantic-ai clai2-bleeding: CLAI2 bleeding

Sdists rebuilt from main for /update on the CLAI bleeding channel.

Code
GitHub: NVIDIA/TensorRT-LLM12d ago

NVIDIA/TensorRT-LLM v1.3.0rc29

Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151

Code
GitHub: pytorch/pytorch11d ago

pytorch/pytorch v2.14.1: PyTorch 2.14.1 Release

This release is meant to fix the following regressions and silent correctness issues:

Code
Repo
Rising AI repositories on GitHub12d ago

sergqwer/strata-nvfp4: Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more. Setup installs our GPTQ quants (h

Strata fork: Qwen3.8-Flash-Next 125B MoE in NVFP4 on one RTX 20-50 card (12 GB+) and 64 GB of RAM or more.

Code
GitHub: ollama/ollama18d ago

ollama/ollama v0.34.4

Qwen 3.8 prompt processing is faster on Apple Silicon.

Code
Hugging Face trending models10d ago

Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw

Infatoshi published the model GLM-5.3-UNCENSORED-EXL3-3.0bpw on Hugging Face.

Weights
GitHub: BerriAI/litellm12h ago

BerriAI/litellm v1.104.3

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm3d ago

BerriAI/litellm v1.104.2

Verify using the pinned commit hash (recommended):

Code
NVIDIA Technical Blog12d ago

AI Native by Design: Lessons Learned from Building NVIDIA TensorRT Model Connect

Parallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT...Parallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT Model Connect is an open source collection of AI model reference implementations in C++, built on top of NVIDIA TensorRT.

GitHub: BerriAI/litellm3d ago

BerriAI/litellm v1.102.4

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm3d ago

BerriAI/litellm v1.101.6

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm4d ago

BerriAI/litellm v1.104.1

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm4d ago

BerriAI/litellm v1.103.4

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm4d ago

BerriAI/litellm v1.102.3

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm4d ago

BerriAI/litellm v1.100.5

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm4d ago

BerriAI/litellm v1.101.5

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm8d ago

BerriAI/litellm v1.104.0

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm7d ago

BerriAI/litellm v1.103.3

Verify using the pinned commit hash (recommended):

Code
GitHub: unslothai/unsloth13d ago

unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library

Run and serve Decision Models like Laya (open-source Jev) locally

Code
GitHub: BerriAI/litellm10d ago

BerriAI/litellm v1.103.2

Verify using the pinned commit hash (recommended):

Code
GitHub: BerriAI/litellm10d ago

BerriAI/litellm v1.101.4

Verify using the pinned commit hash (recommended):

Code
GitHub: vllm-project/vllm19d ago

vllm-project/vllm v0.30.0

This release features 762 commits from 315 contributors (104 new)!

Code
Repo
Rising AI repositories on GitHub12d ago

Dreamer-Toby/STEPQuant: STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization Language: Python.

Code
Repo
Rising AI repositories on GitHub10d ago

StayLameBro/backburner: Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable Language: Python.

Code