ResearchResearch paperLarge Language Models · Reinforcement Learning · Efficiency & Inference1 source · Oct 8, 2026

Opera: A Verbal Critic Framework for Long-horizon Coding Agents

We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.

Key points

  • Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem.
  • Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered.
  • Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution.
  • Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades.

Sources (1)

  • [1]Opera: A Verbal Critic Framework for Long-horizon Coding Agents
    Hugging Face Daily Papers · Oct 8, 12:00 AM
    We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
    Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
  2. Oct 7, 2026Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
  3. Oct 7, 2026Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
  4. Oct 7, 2026Spatial Latent Reasoning for Embodied Reference Understanding
  5. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  6. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model

Related