Policy & safety
Regulation, incidents and safety research.

Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide
An Anthropic AI model provided false information about an unsolved homicide to a Philadelphia Police Department (PPD) tipline, according to a report from 6abc.

OpenAI “rogue” agent activities found on Wikimedia projects
OpenAI “rogue” agent activities found on Wikimedia projects

Investigating unintended model actions in our evaluations and internal use
Investigating unintended model actions in our evaluations and internal use

2026 Usage Policy update
Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers.

When the Safety Test Became the Threat: The Machine That Found Its Own Way Out
OpenAI built a room with no doors – or so it thought.

A big-tent or small-tent AI safety movement?
In recent weeks, two narratives about AI safety have emerged: either AI existential risk is real and imminent, or AI leaders’ and whistleblowers’ claims to that effect are insincere — a “psyop” or hype or a twisted form of regulatory capture.

We tested our own WAF with frontier AI models. Here’s what we found
“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.
Towards safety cases for frontier AI training
Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents

Anthropic is cutting off its internal evaluations from the internet
After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations.
Quoting Victoria Kim
Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said.

GLM-5.3 and the spread of advanced cyber capabilities
Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.