ResearchResearch paperSafety & Alignment · Interpretability1 source · Oct 7, 2026

PairAudit: Guiding Human Review with Graph Tokens under Distribution Shift

Intrusion detectors can confidently misclassify attacks that were not seen during training.

Key points

  • We introduce PairAudit to find overlooked errors and improve review under a fixed budget.
  • Its graph tokens capture prediction patterns across connected nodes.
  • Rather than building another predictor through feature aggregation, PairAudit uses unusual relational patterns to uncover potential errors in existing predictions.
  • Experiments across security tasks show that PairAudit corrects more errors on average than uncertainty-based review, including more errors on unseen attacks.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related