AION
Model

GPT

Also known as: gpt-oss

68stories this week
87last 30 days
95all time

In the model registry

Timeline

  1. Oct 10, 2026 · Opinion / analysis · 1 source
    Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work
    Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
  2. Oct 9, 2026 · Open-source release · 1 source
    crewAIInc/crewAI 1.15.27
    Add deepinfra as an OpenAI-compatible provider
  3. Oct 9, 2026 · Open-source release · 1 source
    pydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)
    <!-- Release notes generated using configuration in .github/release.yml at main -->
  4. Oct 9, 2026 · Opinion / analysis · 1 source
    ICYMI: What landed for AI builders in September 2026
    A recap of the latest Amazon Bedrock, Amazon Bedrock AgentCore, and Strands updates from September 2026
  5. Oct 9, 2026 · Product / feature launch · 1 source
    A new feature for my blog, built using my voice
    I used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment.
  6. Oct 9, 2026 · Opinion / analysis · 1 source
    Asana cuts model costs 76x in browser tests with GPT-6.1 Sol
    Using GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.
  7. Oct 9, 2026 · Opinion / analysis · 1 source
    ttok 1.0
    I released ttok 0.4, ran uv tool upgrade ttok, piped a file into the new version... and realized that it was defaulting to the GPT-4 tokenizer when it should very clearly default to GPT-5/GPT-6 instead!
  8. Oct 8, 2026 · Research paper · 1 source
    SpatialHarness: Test-Time Spatial Scaffolding for Fine Robotic Manipulation
    We introduce SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic manipulation without policy fine-tuning or changes to the physical sensing setup.
  9. Oct 8, 2026 · Research paper · 1 source
    Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
    Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates.
  10. Oct 8, 2026 · Research paper · 1 source
    WOVEN: Weaving Visual World Modeling into Multimodal LLMs
    We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
  11. Oct 8, 2026 · Research paper · 2 sources
    Reasoning-Informed Visual Editing
    To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
  12. Oct 8, 2026 · Research paper · 1 source
    ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
    Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment.
  13. Oct 8, 2026 · Research paper · 1 source
    SteerablePlex: Can We Steer Full-Duplex Models?
    We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order.
  14. Oct 8, 2026 · Research paper · 2 sources
    A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
    We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
  15. Oct 8, 2026 · Research paper · 1 source
    Can Decision Models Understand Stance? Evaluating Jev Against General-Purpose LLMs
    Stance detection requires identifying an author's attitude toward a given target, sometimes based on conversational context.
  16. Oct 8, 2026 · Opinion / analysis · 1 source
    Pollo AI turns creative ideas into campaigns with OpenAI
    With GPT-5.6, GPT-6 Astra, and GPT‐Image‐2.5, Pollo AI helps creators turn bold ideas into detailed images and cinematic video ads.
  17. Oct 8, 2026 · Research paper · 1 source
    RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
    Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.
  18. Oct 8, 2026 · Research paper · 1 source
    Closed-loop evaluation of LLM agents for embedded software development
    We present a benchmark for closed-loop evaluation of embedded coding agents.
  19. Oct 8, 2026 · Research paper · 1 source
    Adversarial Cues in Decision Models Used as Judges: The Role of Request Presentation
    An answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference.
  20. Oct 8, 2026 · Research paper · 1 source
    GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
    We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
  21. Oct 8, 2026 · Research paper · 1 source
    Writing for the Reviewer: Defensive Writing in GPT Models
    Defensive writing grows with GPT version.
  22. Oct 8, 2026 · Research paper · 1 source
    Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild
    We introduce Mine Odyssey, a benchmark for evaluating agentic spatial intelligence using Minecraft reconstructions of real-world locations.
  23. Oct 8, 2026 · Research paper · 1 source
    When Scene Text Hijacks the Scene: Uncovering, Exploiting, and Mitigating Rendered-Text Semantic Leakage in Image Generation Models
    In this work, we study rendered-text semantic leakage, a largely overlooked phenomenon in open-domain text rendering.
  24. Oct 8, 2026 · Research paper · 1 source
    MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
    We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.
  25. Oct 8, 2026 · Research paper · 1 source
    OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
    We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature.
  26. Oct 8, 2026 · Research paper · 1 source
    False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
    To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.
  27. Oct 8, 2026 · Research paper · 1 source
    AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
    To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
  28. Oct 7, 2026 · Research paper · 1 source
    Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models
    Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes.
  29. Oct 7, 2026 · Product / feature launch · 2 sources
    Claude Haiku 5.5
    As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.
  30. Oct 7, 2026 · Research paper · 1 source
    Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models
    We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform.

Often appears with