ResearchResearch paperLarge Language Models · Robotics & Embodied AI · Agents & Tool Use1 source · Oct 6, 2026

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks.

Key points

  • The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing.
  • Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis.
  • Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript.
  • We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use.

Sources (1)

  • [1]Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 01:52 PM
    Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks.
    The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing.

Extractive summary: sentences quoted from the sources.

Related