AION
Research paperReinforcement Learning1 source · Oct 6, 2026

Learning to Report Unsafe Tasks in a Multi-Agent Game

When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task.

Key points

  • Audits can make reporting optimal without ensuring that further training teaches a silent team to report.
  • We study this learning problem in a game where any witness can stop a task by reporting.
  • With $k$ witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared silence probability counts each task's benefit $k$ times at universal silence.
  • For arbitrary policy groups, we give an audit condition sufficient for exact policy-gradient updates to reach universal reporting and, apart from boundary cases, necessary near universal silence.

Sources (1)

  • [1]Learning to Report Unsafe Tasks in a Multi-Agent Game
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:57 PM
    When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task.
    Audits can make reporting optimal without ensuring that further training teaches a silent team to report.

Extractive summary: sentences quoted from the sources.