Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Simon Willison's Weblog2 sources2d ago

llm-openai-decisions 0.1a0

OpenAI released their new Jev-style Decisions API, as previously announced at last week's DevDay.

2 outlets
Video
AI Engineer (YouTube)8h ago

Agent Speedrun — Elizabeth Fuentes Leone & Sandhya Subramani, AWS

A customer service agent looks up an account, checks an order, and handles a refund.

GitHub: openai/codex3d ago

openai/codex rust-v0.162.0: 0.162.0

Add tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. (#50148)

Code
AWS Machine Learning Blog4d ago

Automate remediation post AWS DevOps Agent investigation

In this post, we demonstrate how to use AWS Lambda Durable Functions, a capability of AWS Lambda, Amazon EventBridge, and Amazon Bedrock to create an automated remediation workflow that complements AWS DevOps Agent to complete the issue resolution step.

GitHub: langchain-ai/langchain3d ago

langchain-ai/langchain langchain-openai==1.7.0

chore(deps): bump langgraph-sdk from 0.4.4 to 0.4.6 in /libs/partners/openai (#41096)

Code
Anthropic Research3d ago

Launching an opt-in vulnerability-finding service for open-source software

We’re launching OSS Scanner, an opt-in vulnerability scanner for the open-source ecosystem informed by our experience using Claude to find vulnerabilities during Project Glasswing.

Vendor claim only
GitHub: anthropics/anthropic-sdk-python2d ago

anthropics/anthropic-sdk-python v1.13.0

api: add types for the Chat and Cowork unified analytics metrics

Code
OpenAI News3d ago

How Oracle turns days of work into minutes with ChatGPT and Codex

Across recruiting, engineering, and operations, Oracle turns specialist knowledge into fast, repeatable workflows with ChatGPT Work and Codex.

GitHub: anthropics/claude-code3d ago

anthropics/claude-code v2.1.295

Added onFailure: "block" for command and HTTP hooks: a hook that can't start, times out, or exits with an unexpected code blocks the action instead of letting it through

Code
Ars Technica: AI2d ago

Ukraine’s drones knock out AI data center belonging to "Russia’s Google"

Ukrainian drone strikes have knocked out two of five data centers belonging to the Russian tech giant Yandex.

GitHub: crewAIInc/crewAI2d ago

crewAIInc/crewAI 1.15.27

Add deepinfra as an OpenAI-compatible provider

Code
r/LocalLLaMA (top, daily)20h ago

Qwen3.8 Flash Next fixed my GNOME extension

I love Dash2Dock Lite, but Icedman is always a week or two before updates.

GitHub: BerriAI/litellm14h ago

BerriAI/litellm v1.105.0

Verify using the pinned commit hash (recommended):

Code
GitHub: huggingface/diffusers5d ago

huggingface/diffusers v0.41.0: Diffusers 0.41.0: QwenImage 2.1 pipeline and more

> This release brings Qwen-Image 2.1 to Diffusers, with text-to-image generation, image editing, native transparency, and LoRA training.

Code
r/MachineLearning (top, daily)9h ago

I trained a 414k-parameter transformer to fly a boids flock, then tested whether the rules a probe can read are the ones it uses [P]

I wrote a small boid simulator (12 birds), recorded it flying, and trained a transformer to predict each bird's next move without it knowing about any boid rules.

GitHub: ollama/ollama3d ago

ollama/ollama v0.40.2

Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp.

Code
NVIDIA Technical Blog10d ago

Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills

AI agents are becoming a standard part of development workflows, but general-purpose agents weren't built with specialized infrastructure software such as...AI agents are becoming a standard part of development workflows, but general-purpose agents weren’t built with specialized infrastructure software such as NVIDIA DOCA in mind.

Vendor claim only
GitHub: pydantic/pydantic-ai8d ago

pydantic/pydantic-ai v2.54.0: v2.54.0 (2026-10-02)

<!-- Release notes generated using configuration in .github/release.yml at main -->

Code
Hamel Husain11d ago

Claude’s new auto eval tool

Anthropic released new eval tooling for Claude Code.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks

We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer.

Paper
Latent Space9d ago

[AINews] Pi 1.0, Pi Durable, and AIE NYC

Last call for regular tickets for AI Engineer NYC!

The Register: AI/ML4d ago

COSMIC shuts the door on AI code as GNOME debates letting bug reports in

System76 is banning AI-generated content from contributions to the COSMIC desktop.

Cloudflare Blog: AI11d ago

Cut your AI spend with AI Gateway's Auto Router

Today, we are releasing Cloudflare's Auto Router in public beta, available through AI Gateway.

GitHub: langchain-ai/langgraph5d ago

langchain-ai/langgraph sdk==0.4.6: langgraph-sdk==0.4.6

fix(sdk-py): percent-encode threadid and assistantid in thread stream requests (#9213)

Code
GitHub: openai/openai-agents-python9d ago

openai/openai-agents-python v0.23.1

Release readiness review (v0.23.0 -> TARGET 0a5a08b21cd44808d73e8a9e732212dd33b3690d)

Code
Last Week in AI8d ago

LWiAI Podcast #258 - Opus 5.5, Sol and Luna, Muse, DeepSeek-V4.1-Flash, Xi

Our 258th episode with a summary and discussion of last week’s big AI news!

Philipp Schmid13d ago

Credentials API for Gemini Managed Agents

Store secrets once with the Credentials API, reference them by ID, and let the egress proxy inject them on the wire so they never enter the agent sandbox.

Hugging Face Blog13d ago

Welcome RL Environments to the hub

An environment gives an agent a task, responds to its actions with observations, and scores the outcome.

Vendor claim only
GitHub: NVIDIA/TensorRT-LLM12d ago

NVIDIA/TensorRT-LLM v1.3.0rc29

Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151

Code
GitHub: modelcontextprotocol/python-sdk9d ago

modelcontextprotocol/python-sdk v2.3.0

A few things behave differently, so skim these first:

Code
GitHub: unslothai/unsloth10d ago

unslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UX

This release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.

Code
GitHub: ray-project/ray9d ago

ray-project/ray ray-2.59.0: Ray-2.59.0

Ray Data LLM & Ray Serve LLM are GA/Stable: the LLM APIs graduate to general availability this release (\#65194), alongside an upgrade to vLLM 0.27.0 (\#65351).

Code
Ollama Blog12d ago

Ollama now supports Jev-style decision models

Ollama now supports decision models, based on TypeSafe's Jev API for fast, typed decisions.

xAI News13d ago

Team Bots: shared AI teammates that learn as they work

Today we’re launching Team Bots, Grok Bots that work and learn alongside your team.

GitHub: vllm-project/vllm19d ago

vllm-project/vllm v0.30.0

This release features 762 commits from 315 contributors (104 new)!

Code
Google Blog (AI)23d ago

Co-creating the future of fashion with Google

Google worked side-by-side with designers Jane Wade and Sergio Hudson to custom-design Google Flow tools to prep for NYFW.

GitHub: openai/codex2d ago

openai/codex rust-v0.162.1: 0.162.1

Fixed a TUI crash when asynchronous questions contain multiple lines, preserving line breaks and complete hyperlink destinations. (#51866)

Code
AWS Machine Learning Blog5d ago

Manage Amazon SageMaker HyperPod Spaces directly from SageMaker Studio

We recently introduced the ability to create and manage Amazon SageMaker Spaces on Amazon SageMaker HyperPod EKS clusters directly from the Amazon SageMaker Studio UI.

AWS Machine Learning Blog2d ago

ICYMI: What landed for AI builders in September 2026

A recap of the latest Amazon Bedrock, Amazon Bedrock AgentCore, and Strands updates from September 2026

Simon Willison's Weblog2d ago

ttok 1.0

I released ttok 0.4, ran uv tool upgrade ttok, piped a file into the new version... and realized that it was defaulting to the GPT-4 tokenizer when it should very clearly default to GPT-5/GPT-6 instead!

Simon Willison's Weblog5d ago

llm-mistral 0.16

Adds support for reasoning models, such as the newly released Mistral Large 4.

r/LocalLLaMA (top, daily)1d ago

Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work

Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).

OpenAI News5d ago

Atlassian and OpenAI expand partnership to turn enterprise knowledge into action

Atlassian and OpenAI are expanding their partnership to connect frontier models with enterprise knowledge and help teams plan, build, and deliver work.

OpenAI News3d ago

LegalOn halves Codex costs while maintaining development speed

LegalOn cut estimated daily Codex costs by 65% while maintaining development speed.

AWS Machine Learning Blog3d ago

Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod

Multiple teams within the same company increasingly need shared access to expensive GPU clusters for their generative AI operations, while maintaining isolation boundaries, resource fairness, and operational independence.

r/LocalLLaMA (top, daily)1d ago

[Model] Support MiniCPM-V 4.7 by tc-mb · Pull Request #29416 · ggml-org/llama.cpp

Let me remind you that MiniCPM-V-4.7-35B-A3B was spotted on r/LocalLLaMA a few days ago (but the model was later hidden on HF).

Simon Willison's Weblog23h ago

Python 3.15.0 added to actions/python-versions

Now that this has landed, you can add "3.15" to a GitHub Actions testing matrix to run tests against the new Python 3.15.0 release.

GitHub: anthropics/anthropic-sdk-python3d ago

anthropics/anthropic-sdk-python v1.12.1

docs: note that listing Claude Console spend limits is in early access

Code