Learning to Report Unsafe Tasks in a Multi-Agent Game
When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task.
Key points
- Audits can make reporting optimal without ensuring that further training teaches a silent team to report.
- We study this learning problem in a game where any witness can stop a task by reporting.
- With $k$ witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared silence probability counts each task's benefit $k$ times at universal silence.
- For arbitrary policy groups, we give an audit condition sufficient for exact policy-gradient updates to reach universal reporting and, apart from boundary cases, necessary near universal silence.
Sources (1)
- [1]Learning to Report Unsafe Tasks in a Multi-Agent GamearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:57 PM
When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task.
Audits can make reporting optimal without ensuring that further training teaches a silent team to report.
Extractive summary: sentences quoted from the sources.