AION
Research paperSafety & Alignment · Agents & Tool Use1 source · Oct 7, 2026

From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage

Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert.

Key points

  • Security operations centers (SOCs) must triage large volumes of alerts, most of which are benign, while missed attacks can remain uninvestigated.
  • To support this study, we build ALERT-BENCH, an interactive benchmark that replays enterprise telemetry through a live SIEM and requires each system to retrieve evidence.
  • Based on these findings, we further design AIDA (Adversarial Investigation and Dialectical Analysis), a multi-agent framework that requires an explicit proposed decision before independent challenge and stronger evidentiary requirements before dismissal.
  • These results show that structuring evidence retrieval and decision review can substantially improve agentic SOC triage.

Sources (1)

  • [1]From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:20 AM
    Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert.
    Security operations centers (SOCs) must triage large volumes of alerts, most of which are benign, while missed attacks can remain uninvestigated.

Extractive summary: sentences quoted from the sources.