Policy & safety

Regulation, incidents and safety research.

TechCrunch: AI4 sources2d ago

Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide

An Anthropic AI model provided false information about an unsolved homicide to a Philadelphia Police Department (PPD) tipline, according to a report from 6abc.

2 outlets
Simon Willison's Weblog2 sources4d ago

OpenAI “rogue” agent activities found on Wikimedia projects

OpenAI “rogue” agent activities found on Wikimedia projects

2 outlets
Anthropic Research2d ago

Investigating unintended model actions in our evaluations and internal use

Investigating unintended model actions in our evaluations and internal use

Anthropic News3d ago

2026 Usage Policy update

Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers.

MarkTechPost1d ago

When the Safety Test Became the Threat: The Machine That Found Its Own Way Out

OpenAI built a room with no doors – or so it thought.

AWS Machine Learning Blog4d ago

Automate remediation post AWS DevOps Agent investigation

In this post, we demonstrate how to use AWS Lambda Durable Functions, a capability of AWS Lambda, Amazon EventBridge, and Amazon Bedrock to create an automated remediation workflow that complements AWS DevOps Agent to complete the issue resolution step.

Ars Technica: AI2 sources2d ago

AI disqualification yields new Nikon Small World in Motion winner

Last month we covered the winner of Nikon's Small World in Motion video: Ning Xu of Tsinghua University in China, whose video captured tiny cilia beating in the airways of a child with a rare respiratory disorder.

2 outlets
The Register: AI/ML4d ago

COSMIC shuts the door on AI code as GNOME debates letting bug reports in

System76 is banning AI-generated content from contributions to the COSMIC desktop.

The Verge: AI1d ago

DistroKid has been quietly taking down songs in response to UMG lawsuit

Artists are taking to social media to complain that DistroKid has unceremoniously removed their work without notice.

Exponential View (Azeem Azhar)7d ago

🔮 The transition is hiding in plain sight #604

We are continuing to push hard on untangling this data and more over at our AI Investment Brief by Exponential View.

AI Snake Oil10d ago

A big-tent or small-tent AI safety movement?

In recent weeks, two narratives about AI safety have emerged: either AI existential risk is real and imminent, or AI leaders’ and whistleblowers’ claims to that effect are insincere — a “psyop” or hype or a twisted form of regulatory capture.

WIRED: AI2d ago

Even ‘Law & Order’ Is Terrified of AI

In its 26th season premiere, the procedural legal drama paints a damning picture of power-mad AI CEOs.

Cloudflare Blog: AI9d ago

Building for good: How civil society organizations are automating on Cloudflare

The goal of Cloudflare Impact is to help ensure that non-profit organizations are among the first to benefit.

Product Hunt: AI launches5d ago

Regunow

Turn regulatory research into actionable compliance audits

OpenAI News13d ago

Towards safety cases for frontier AI training

Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents

TechCrunch: AI2 sources1d ago

Anthropic is cutting off its internal evaluations from the internet

After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations.

2 outlets
Simon Willison's Weblog4d ago

Quoting Victoria Kim

Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said.

Exponential View (Azeem Azhar)6d ago

📈 Monday data: More AI, more justice?

Latest in our AI Investment Brief: An update on AI revenues + what’s happening with AI spending.

Anthropic Research12d ago

GLM-5.3 and the spread of advanced cyber capabilities

Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.

Ars Technica: AI6d ago

AI glasses face their first major government crackdown

Norway has become the first major country to propose a temporary ban on the use of AI glasses in selected public places amid growing privacy concerns over wearable technology.

Ars Technica: AI4d ago

Fraudster jailed for using 10K bots and AI songs to outstream Taylor Swift

After pleading guilty, a 54-year-old North Carolina man, Michael Smith, was sentenced to 18 months in prison for using artificial intelligence to generate songs for a scheme that stole millions from music streaming platforms.

Cloudflare Blog: AI12d ago

We tested our own WAF with frontier AI models. Here’s what we found

“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.