Policy & safety

Regulation, incidents and safety research.

TechCrunch: AI4 sources2d ago

Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide

An Anthropic AI model provided false information about an unsolved homicide to a Philadelphia Police Department (PPD) tipline, according to a report from 6abc.

2 outlets
MarkTechPost1d ago

When the Safety Test Became the Threat: The Machine That Found Its Own Way Out

OpenAI built a room with no doors – or so it thought.

Anthropic Research2d ago

Investigating unintended model actions in our evaluations and internal use

Investigating unintended model actions in our evaluations and internal use

Anthropic News3d ago

2026 Usage Policy update

Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers.

Simon Willison's Weblog2 sources4d ago

OpenAI “rogue” agent activities found on Wikimedia projects

OpenAI “rogue” agent activities found on Wikimedia projects

2 outlets
TechCrunch: AI2 sources1d ago

Anthropic is cutting off its internal evaluations from the internet

After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations.

2 outlets
Simon Willison's Weblog4d ago

Quoting Victoria Kim

Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said.