Hidden gems
Fast-rising repos, papers the community loves, and posts from trusted independent researchers, before the press picks them up.
From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation
We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation.

Investigating unintended model actions in our evaluations and internal use
Investigating unintended model actions in our evaluations and internal use
Expanding the Cyber Verification Program
We’re launching a new, expanded version of our Cyber Verification Program (CVP), which makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals.
MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation.

Introducing the Anthropic Cyber Mission
Today we’re launching the Anthropic Cyber Mission, a long-term commitment to securing the systems everyone depends on.
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
Building on our commitment to American scientific discovery
Building on our commitment to American scientific discovery
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

2026 Usage Policy update
Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers.
SuperNav: An Agentic Navigation System for Any Task in Any Scene
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.

Launching an opt-in vulnerability-finding service for open-source software
We’re launching OSS Scanner, an opt-in vulnerability scanner for the open-source ecosystem informed by our experience using Claude to find vulnerabilities during Project Glasswing.
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments.
EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2: an open, lightweight multimodal embedding model
AgentGarten: Code Worlds for Evolving Agents
We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.
Atlassian and OpenAI expand partnership to turn enterprise knowledge into action
Atlassian and OpenAI are expanding their partnership to connect frontier models with enterprise knowledge and help teams plan, build, and deliver work.
TokenRouter: Efficient Serving System for Token-Level LLM Routing
Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving.
Pollo AI turns creative ideas into campaigns with OpenAI
With GPT-5.6, GPT-6 Astra, and GPT‐Image‐2.5, Pollo AI helps creators turn bold ideas into detailed images and cinematic video ads.
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
Sophos cuts threat investigation time by 96% with OpenAI Daybreak
Discover how Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96% and automate 52% of MDR cases while preserving human oversight.
OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.
Asana cuts model costs 76x in browser tests with GPT-6.1 Sol
Using GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
How Jump Trading is scaling quant research with ChatGPT
Jump Trading uses OpenAI to expand quantitative research.
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory.
Unlocking Earth AI’s planetary geospatial foundation models for global public health
In our latest work, we present five partner-driven case studies demonstrating how this model exemplifies the planetary geospatial foundation model paradigm for global public health.
Opera: A Verbal Critic Framework for Long-horizon Coding Agents
We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
Radisson Hotel Group brings hotel discovery into ChatGPT
Radisson partnered with Accenture to build a ChatGPT plugin using OpenAI technology, helping travelers find, compare, and book hotels while planning their trips.
U-Space: Uncovering When and Why Uncertainty Arises in Language Models
Large language models are informing decisions with ever-higher stakes.
How Oracle turns days of work into minutes with ChatGPT and Codex
Across recruiting, engineering, and operations, Oracle turns specialist knowledge into fast, repeatable workflows with ChatGPT Work and Codex.
Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.
Quoting Victoria Kim
Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said.
SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
Quoting Felix Rieseberg
The "old" version of Cowork runs model inference in the cloud, executing tool calls in an Anthropic-provided VM we shipped to your computer.
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.

Anthropic Subscriptions Offer 5x+ More Value Than OpenAI
Subscription plans are still the primary way consumers and small businesses pay for AI.
REMORY: Learning Residual Memory for Context Compaction
We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens.
Scaling Decision Optimization to 100 Million Variables and Beyond with mPDLP in NVIDIA cuOpt
NVIDIA cuOpt GPU-accelerated decision optimization can already deliver speedups of more than 10x over CPU...
OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
A new feature for my blog, built using my voice
I used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment.
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.