Terminal-Bench
7stories this week
8last 30 days
9all time
Timeline
- Oct 8, 2026 · Research paper · 2 sourcesREMORY: Learning Residual Memory for Context CompactionWe introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens.
- Oct 8, 2026 · Research paper · 1 sourceOpera: A Verbal Critic Framework for Long-horizon Coding AgentsWe present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
- Oct 7, 2026 · Research paper · 1 sourceCoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-EvolutionWithin this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh.
- Oct 7, 2026 · Research paper · 1 sourceCache the Encoder Within:Compact, Reusable Memory across LLM QueriesRepeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.
- Oct 6, 2026 · Research paper · 1 sourceFreeEvolve: Learning to Evolve Beyond Fixed LoopsAgent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops.
- Oct 6, 2026 · Research paper · 1 sourceHarness Engineering for Software Engineering via Modular Executable Dev-PrimitivesBuilding on Dev-Primitives, we propose HERMES, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised.
- Oct 5, 2026 · Product / feature launch · 1 sourceIntroducing GLM 5.3 on Amazon BedrockGLM 5.3 from Z.ai (Zhipu AI) is now available on Amazon Bedrock.
- Sep 28, 2026 · Product / feature launch · 1 sourceWelcome RL Environments to the hubAn environment gives an agent a task, responds to its actions with observations, and scores the outcome.
- May 15, 2026 · Product / feature launch · 1 sourceGemini 3.5: frontier intelligence with actionGemini 3.5: frontier intelligence with action