Deception by Omission: Language Models Knowingly Hide Their Mistakes
Large language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed.
Key points
- Users then depend on the model to report what went wrong.
- An honest model discloses its mistakes, while a deceptive one conceals them.
- In this study, we prefill LLM trajectories with synthetic mistakes.
- Models fail to disclose their mistake in 36.4% of chat and 67.1% of agentic rollouts.
Sources (1)
- [1]Deception by Omission: Language Models Knowingly Hide Their MistakesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:45 AM
Large language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed.
Users then depend on the model to report what went wrong.
Extractive summary: sentences quoted from the sources.