Codex
22stories this week
26last 30 days
26all time
Timeline
- Oct 11, 2026 · Open-source release · 1 sourceBerriAI/litellm v1.105.0Verify using the pinned commit hash (recommended):
- Oct 10, 2026 · Opinion / analysis · 1 sourceEngineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering workModels: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
- Oct 10, 2026 · Opinion / analysis · 1 sourceAre .ipynb notebooks already outdated in the agentic era? [D]Back then, Jupyter Notebooks were a perfect fit for the classical DS pipeline: EDA -> data prep -> fit -> eval -> tune -> save model artefact and notebook.
- Oct 9, 2026 · Open-source release · 1 sourcepydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)<!-- Release notes generated using configuration in .github/release.yml at main -->
- Oct 9, 2026 · Product / feature launch · 1 sourceA new feature for my blog, built using my voiceI used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment.
- Oct 9, 2026 · Opinion / analysis · 1 sourceAsana cuts model costs 76x in browser tests with GPT-6.1 SolUsing GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.
- Oct 8, 2026 · Opinion / analysis · 1 sourcePay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore paymentsIn this post, we look at how Incarna used AgentCore payments to let its agents pay BlockRun for model inference one request at a time.
- Oct 8, 2026 · Opinion / analysis · 1 sourceHow Oracle turns days of work into minutes with ChatGPT and CodexAcross recruiting, engineering, and operations, Oracle turns specialist knowledge into fast, repeatable workflows with ChatGPT Work and Codex.
- Oct 8, 2026 · Research paper · 2 sourcesA Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box OptimizationWe therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
- Oct 8, 2026 · Opinion / analysis · 1 sourceLegalOn halves Codex costs while maintaining development speedLegalOn cut estimated daily Codex costs by 65% while maintaining development speed.
- Oct 8, 2026 · Research paper · 1 sourceSWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn InteractionTo address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.
- Oct 8, 2026 · Research paper · 1 sourceSkill Constellations: Tracing the Supply Chain of Agent Skills on GitHubAgent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user.
- Oct 7, 2026 · Research paper · 1 sourceCross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection StudyWe describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence.
- Oct 7, 2026 · Product / feature launch · 2 sourcesClaude Haiku 5.5As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.
- Oct 7, 2026 · Product / feature launch · 1 sourceAgent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real HarnessesHarnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.
- Oct 7, 2026 · Research paper · 1 sourceQuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced AgentsHere we present QuSema, an autonomous testing agent for finding silent bugs in quantum libraries.
- Oct 7, 2026 · Research paper · 1 sourceAgentTime: Can Agents Estimate and Control Their Own Runtime?We present AgentTime, a benchmark for testing whether agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward.
- Oct 6, 2026 · Tutorial / explainer · 1 sourceUsing Parseable with Datasette for OpenTelemetry tracesTIL: Using Parseable with Datasette for OpenTelemetry traces
- Oct 6, 2026 · Research paper · 2 sourcesCan AI Agents Make Open-Ended Scientific Discovery? Evidence from StationRecent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear.
- Oct 6, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.903-beta: New Browser + Voice CloningThis release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
- Oct 5, 2026 · Product / feature launch · 1 sourceNew agent skill: Amazon SageMaker optimized generative AI inference for your coding agentToday, Amazon SageMaker AI optimized generative AI inference introduces the aws-ai-ml skill, available through the Agent Toolkit for AWS.
- Oct 4, 2026 · Tutorial / explainer · 1 sourceQwen3.8 27B addition in wordsResearch: Qwen3.8 27B addition in words
- Oct 2, 2026 · Product / feature launch · 1 sourceChatham scales its capital markets expertise with OpenAIChatham Financial uses Codex and GPT-5.6 to build technology and redesign workflows, cutting trade validation from 30 minutes to under 4.
- Sep 29, 2026 · Product / feature launch · 1 sourceDevDay 2026 RecapExplore more than 20 announcements from OpenAI DevDay 2026, including GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.
- Sep 28, 2026 · Opinion / analysis · 1 sourceAre you a Codex Original?We’re collecting real stories of builders, tinkerers, researchers, and creators who are using Codex to do incredible things.
- Sep 23, 2026 · Open-source release · 1 sourceollama/ollama v0.34.4Qwen 3.8 prompt processing is faster on Apple Silicon.