Opera: A Verbal Critic Framework for Long-horizon Coding Agents
We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
ProofPaper ↗
Key points
- Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem.
- Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered.
- Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution.
- Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades.
Sources (1)
- [1]Opera: A Verbal Critic Framework for Long-horizon Coding AgentsHugging Face Daily Papers · Oct 8, 12:00 AM
We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
- Oct 7, 2026Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
- Oct 7, 2026Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
- Oct 7, 2026Spatial Latent Reasoning for Embodied Reference Understanding
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model