Policy & safety
Regulation, incidents and safety research.

TechCrunch: AI4 sources2d ago
Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide
An Anthropic AI model provided false information about an unsolved homicide to a Philadelphia Police Department (PPD) tipline, according to a report from 6abc.
2 outlets

Simon Willison's Weblog2 sources4d ago
OpenAI “rogue” agent activities found on Wikimedia projects
OpenAI “rogue” agent activities found on Wikimedia projects
2 outlets

MarkTechPost1d ago
When the Safety Test Became the Threat: The Machine That Found Its Own Way Out
OpenAI built a room with no doors – or so it thought.
The Register: AI/ML4d ago
COSMIC shuts the door on AI code as GNOME debates letting bug reports in
System76 is banning AI-generated content from contributions to the COSMIC desktop.

Anthropic Research12d ago
GLM-5.3 and the spread of advanced cyber capabilities
Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.

TechCrunch: AI2 sources1d ago
Anthropic is cutting off its internal evaluations from the internet
After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations.
2 outlets