GPT
Also known as: gpt-oss
68stories this week
87last 30 days
95all time
In the model registry
| Model | Params | Context | Released |
|---|---|---|---|
| nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 | 27B | – | Oct 4, 2026 |
| openai/gpt-oss-safeguard-20b | 21.5B | 128K | Sep 18, 2025 |
| openai/gpt-oss-safeguard-120b | 120B | 128K | Sep 18, 2025 |
| openai/gpt-oss-20b | 20.9B | 128K | Aug 4, 2025 |
| openai/gpt-oss-120b | 117B | 128K | Aug 4, 2025 |
Timeline
- Oct 10, 2026 · Opinion / analysis · 1 sourceEngineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering workModels: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
- Oct 9, 2026 · Open-source release · 1 sourcecrewAIInc/crewAI 1.15.27Add deepinfra as an OpenAI-compatible provider
- Oct 9, 2026 · Open-source release · 1 sourcepydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)<!-- Release notes generated using configuration in .github/release.yml at main -->
- Oct 9, 2026 · Opinion / analysis · 1 sourceICYMI: What landed for AI builders in September 2026A recap of the latest Amazon Bedrock, Amazon Bedrock AgentCore, and Strands updates from September 2026
- Oct 9, 2026 · Product / feature launch · 1 sourceA new feature for my blog, built using my voiceI used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment.
- Oct 9, 2026 · Opinion / analysis · 1 sourceAsana cuts model costs 76x in browser tests with GPT-6.1 SolUsing GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.
- Oct 9, 2026 · Opinion / analysis · 1 sourcettok 1.0I released ttok 0.4, ran uv tool upgrade ttok, piped a file into the new version... and realized that it was defaulting to the GPT-4 tokenizer when it should very clearly default to GPT-5/GPT-6 instead!
- Oct 8, 2026 · Research paper · 1 sourceSpatialHarness: Test-Time Spatial Scaffolding for Fine Robotic ManipulationWe introduce SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic manipulation without policy fine-tuning or changes to the physical sensing setup.
- Oct 8, 2026 · Research paper · 1 sourceRounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State QuantizationQuantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates.
- Oct 8, 2026 · Research paper · 1 sourceWOVEN: Weaving Visual World Modeling into Multimodal LLMsWe therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
- Oct 8, 2026 · Research paper · 2 sourcesReasoning-Informed Visual EditingTo study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
- Oct 8, 2026 · Research paper · 1 sourceReSI: Recursive Safety Improvement toward Resistant and Resilient AIRecursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment.
- Oct 8, 2026 · Research paper · 1 sourceSteerablePlex: Can We Steer Full-Duplex Models?We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order.
- Oct 8, 2026 · Research paper · 2 sourcesA Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box OptimizationWe therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
- Oct 8, 2026 · Research paper · 1 sourceCan Decision Models Understand Stance? Evaluating Jev Against General-Purpose LLMsStance detection requires identifying an author's attitude toward a given target, sometimes based on conversational context.
- Oct 8, 2026 · Opinion / analysis · 1 sourcePollo AI turns creative ideas into campaigns with OpenAIWith GPT-5.6, GPT-6 Astra, and GPT‐Image‐2.5, Pollo AI helps creators turn bold ideas into detailed images and cinematic video ads.
- Oct 8, 2026 · Research paper · 1 sourceRouterInterp: Understanding Superposed Specialisation in Mixture of Experts RoutingLeveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.
- Oct 8, 2026 · Research paper · 1 sourceClosed-loop evaluation of LLM agents for embedded software developmentWe present a benchmark for closed-loop evaluation of embedded coding agents.
- Oct 8, 2026 · Research paper · 1 sourceAdversarial Cues in Decision Models Used as Judges: The Role of Request PresentationAn answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference.
- Oct 8, 2026 · Research paper · 1 sourceGroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQAWe present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
- Oct 8, 2026 · Research paper · 1 sourceWriting for the Reviewer: Defensive Writing in GPT ModelsDefensive writing grows with GPT version.
- Oct 8, 2026 · Research paper · 1 sourceMine Odyssey: Benchmarking Spatial Agentic Intelligence in the WildWe introduce Mine Odyssey, a benchmark for evaluating agentic spatial intelligence using Minecraft reconstructions of real-world locations.
- Oct 8, 2026 · Research paper · 1 sourceWhen Scene Text Hijacks the Scene: Uncovering, Exploiting, and Mitigating Rendered-Text Semantic Leakage in Image Generation ModelsIn this work, we study rendered-text semantic leakage, a largely overlooked phenomenon in open-domain text rendering.
- Oct 8, 2026 · Research paper · 1 sourceMiniVer-V: Identifying Minimal Sufficient Evidence for Short Video VerificationWe introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.
- Oct 8, 2026 · Research paper · 1 sourceOpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical SciencesWe introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature.
- Oct 8, 2026 · Research paper · 1 sourceFalse Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image GeneratorsTo fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.
- Oct 8, 2026 · Research paper · 1 sourceAgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use TasksTo this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
- Oct 7, 2026 · Research paper · 1 sourceOmni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language ModelsUnified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes.
- Oct 7, 2026 · Product / feature launch · 2 sourcesClaude Haiku 5.5As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.
- Oct 7, 2026 · Research paper · 1 sourceGrammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier ModelsWe close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform.