Open source
Libraries, tools and repos: releases you depend on and projects gaining stars.

llm-openai-decisions 0.1a0
OpenAI released their new Jev-style Decisions API, as previously announced at last week's DevDay.
pydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)
<!-- Release notes generated using configuration in .github/release.yml at main -->
langchain-ai/langchain langchain-openai==1.7.0
chore(deps): bump langgraph-sdk from 0.4.4 to 0.4.6 in /libs/partners/openai (#41096)
Asana cuts model costs 76x in browser tests with GPT-6.1 Sol
Using GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.
huggingface/trl v1.15.0
SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
anthropics/anthropic-sdk-python v1.13.0
api: add types for the Chat and Cowork unified analytics metrics
ausboss/Qwen-Image-2.1-Outfit-Swap-Consistency-LoRA
ausboss published the model Qwen-Image-2.1-Outfit-Swap-Consistency-LoRA on Hugging Face.
huggingface/transformers v5.19.0: Release v5.19.0
EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
crewAIInc/crewAI 1.15.27
Add deepinfra as an OpenAI-compatible provider
VideoParameter Golf with AutoResearch — Vayum Arora, Zhengyao Jiang, Dixing Xu & Dhruv Srikanth, Weco AI
Zhengyao Jiang introduces autoresearch as repeated proposals and evaluations, and Dixing Xu explains the team's Aiden system and its contributions to OpenAI's Parameter Golf challenge.
Converting dense models into Mixture-of-Experts
For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
BerriAI/litellm v1.105.0
Verify using the pinned commit hash (recommended):
huggingface/diffusers v0.41.0: Diffusers 0.41.0: QwenImage 2.1 pipeline and more
> This release brings Qwen-Image 2.1 to Diffusers, with text-to-image generation, image editing, native transparency, and LoRA training.
open-webui/open-webui v0.12.0
Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.
Faster Scientific Image Analysis with NVIDIA cuPhoton
Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can...Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can process it to support timely decisions.
unslothai/unsloth v0.1.905-beta: Sandboxing is here!
We're introducing Windows, Mac and Linux sandboxing in Unsloth!
vllm-project/vllm v0.31.0
Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
anthropics/claude-code v2.1.294
Fixed prompt and agent hooks written as instructions (such as "Block commands that...") allowing what they should block
ollama/ollama v0.40.2
Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp.
browser-use/browser-use 0.13.11
This release includes Browser Use toolsets for Claude, available at browseruse.integrations.toolsetsforclaude.
[AINews] Pi 1.0, Pi Durable, and AIE NYC
Last call for regular tickets for AI Engineer NYC!
langchain-ai/langgraph cli==0.4.33: langgraph-cli==0.4.33
feat(cli): add 'langgraph deploy listeners list' (#9221)
GTKottman/mortiflix-oss: A motion design studio on your own machine: Claude makes the video step by step, you approve every stage. Bring your own Claude Code or API
A motion design studio on your own machine: Claude makes the video step by step, you approve every stage.
NVIDIA/TensorRT-LLM v1.3.0rc29
Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
modelcontextprotocol/python-sdk v2.3.0
A few things behave differently, so skim these first:
sgl-project/sglang v0.5.21
| Model | Type | Cookbook |
ray-project/ray ray-2.59.0: Ray-2.59.0
Ray Data LLM & Ray Serve LLM are GA/Stable: the LLM APIs graduate to general availability this release (\#65194), alongside an upgrade to vLLM 0.27.0 (\#65351).
pytorch/pytorch v2.14.1: PyTorch 2.14.1 Release
This release is meant to fix the following regressions and silent correctness issues:
huggingface/peft v0.21.2
This is a PEFT release fixes an issue that prevented encoder-decoder models to work when using Transformers ≥ 5.18.0.
speridlabs/iris-3b
speridlabs published the model iris-3b on Hugging Face.
LegalOn halves Codex costs while maintaining development speed
LegalOn cut estimated daily Codex costs by 65% while maintaining development speed.
How DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software Stack
GPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services.
Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit
Mia-AiLab published the model GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit on Hugging Face.
Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw
Infatoshi published the model GLM-5.3-UNCENSORED-EXL3-3.0bpw on Hugging Face.
unslothai/unsloth v0.1.904-beta: Train your own Decision model
Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.
langchain-ai/langchain langchain-huggingface==1.2.3
chore(deps): bump langgraph-sdk from 0.4.4 to 0.4.6 in /libs/partners/huggingface (#41098)
huggingface/trl v1.14.2
Patch release fixing two cases of silently wrong training and three crashes.
anthropics/anthropic-sdk-python v1.12.1
docs: note that listing Claude Console spend limits is in early access
langchain-ai/langchain langchain==1.4.4
chore(deps): bump langgraph-sdk from 0.4.2 to 0.4.4 in /libs/langchainv1 (#41074)
langchain-ai/langchain langchain-core==1.6.9
feat(core): accept a callable in withretry (#41158)
pydantic/pydantic-ai v2.54.0: v2.54.0 (2026-10-02)
<!-- Release notes generated using configuration in .github/release.yml at main -->
langchain-ai/langchain langchain-core==1.6.8
fix(core): harden scoped and compatible IPv6 SSRF checks (#41151)
canberkkkkkk/ema-lightning
canberkkkkkk published the model ema-lightning on Hugging Face.
Cloudflare/clef-flash
Cloudflare published the model clef-flash on Hugging Face.
ttok 0.4
ttok is my CLI tool for counting tokens, using OpenAI's open source tiktoken library.
langchain-ai/langchain langchain-fireworks==1.7.1
chore(deps): bump langgraph-sdk from 0.4.4 to 0.4.6 in /libs/partners/fireworks (#41100)
Cactus-Compute/whistle
Cactus-Compute published the model whistle on Hugging Face.

Can a Cloud-Native Harness Make Agents Reliable Beyond the Desktop?
Since that’s easiest to implement locally, many coding agents started out as terminal tools and then desktop apps (Claude Code famously began as a CLI tool).