Explore
Everything AION read, in seven sections. Pick one, a topic or a time window.

Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide
An Anthropic AI model provided false information about an unsolved homicide to a Philadelphia Police Department (PPD) tipline, according to a report from 6abc.

Satya Nadella says we should assume all AI models are ‘compromised’
In a lengthy post on X, Microsoft's CEO laid out his views on the dangers posed by highly advanced AI models and how to confront those risks.

OpenAI “rogue” agent activities found on Wikimedia projects
OpenAI “rogue” agent activities found on Wikimedia projects

Google rolls out improved SynthID AI content detector, now available globally
Every piece of AI content from Google's Gemini models has a hidden SynthID label, which makes it very difficult to pass the content off as authentic.
Our approach to EU text provenance rules
How OpenAI is approaching text watermarking under EU rules.

Investigating unintended model actions in our evaluations and internal use
Investigating unintended model actions in our evaluations and internal use

Introducing the Anthropic Cyber Mission
Today we’re launching the Anthropic Cyber Mission, a long-term commitment to securing the systems everyone depends on.

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context
OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research.
pydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)
<!-- Release notes generated using configuration in .github/release.yml at main -->
I Made Terrible Games With Google’s AI Playground
A long day’s haul in the video game slop mines.

[AINews] not much happened today
More AI safety intrigue in the alignment below.

OpenAI says a misaligned model deliberately destroyed its own environment hoping for a fresh start with better data
OpenAI has documented new cases of misaligned model behavior.
VideoAI Security Engineer Foundations + Certificate — Micah Silverman, Snyk
Micah Silverman uses that capture the flag exercise to connect AI risks with familiar application security controls.
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution.
From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents
In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope.
BerriAI/litellm v1.105.0
Verify using the pinned commit hash (recommended):

Can AI automate AI R&D yet?
Existing evidence shows that AI can perform software engineering tasks relevant to AI research, dataset creation, open-ended optimization of defined metrics (autoresearch), and more.
Toward provably private learning from federated data
In 2017, Google introduced Federated Learning (FL) a machine learning technique that trains models across decentralized, private data.
anthropics/claude-code v2.1.294
Fixed prompt and agent hooks written as instructions (such as "Block commands that...") allowing what they should block
What AI gets wrong and what failure teaches us
Jennifer Neville is a partner research manager at Microsoft who’s built a career around understanding and advancing AI for real-world use, and much like the human-AI interactions she’s been studying, her early-career path was multiturn: math, then physics; cognitive science, then work; and finally computer science—despite her best efforts to avoid the field.

Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots
Anthropic and OpenAI race to release smarter and cheaper models
Validate AI Factory Changes with Digital Twins and AI Agents
AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration...AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration services, security controls, and a rapidly changing software stack.
huggingface/trl v1.14.2
Patch release fixing two cases of silently wrong training and three crashes.
Responsible AI governance: How AWS positions customers to align with ISO/IEC 42005:2025
With generative AI adoption moving faster than the personal computer or the internet and global AI-related investment in 2025 representing $581.69 billion, organizations must position their workforce to use AI to power their operations while employing it responsibly.

Attempts to Keep Humans in the AI Loop May Actually Push Them Out
A crucial safeguard against AI agents going rogue—keeping humans in the loop to review and approve their decisions—will fail unless designers and users change their current practices, a trio of leading AI ethics researchers argue.
langchain-ai/langgraph sdk==0.4.6: langgraph-sdk==0.4.6
fix(sdk-py): percent-encode threadid and assistantid in thread stream requests (#9213)
Secret protection must scale with software
Today, one in three pull requests on GitHub involves an AI agent.
The Cyber Risk Discourse is Broken
Western voices saying open weight models are necessary for defense and banning them will make the world less safe, occupied by AI risk moderates to different extremes.
VideoOpenAI Security: Controlling Models is Now ‘Hell’
A cracked cipher, an OpenAI security warning, Gemini 4 Argon, RSI paper (co-authored by a who’s who of AI), Lab White House commitments, new hacks emerging, ‘deep personas’, biology Kasparov competitions, and so much more, ending with an epic Opus outro.
People really hate AI, so why can’t they get enough?
Over the summer I talked to the CEO of Springboards, a startup building an LLM that’s designed to come up with a wider variety of responses than its mainstream rivals do.
Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Our latest paper, ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (read it on Hugging Face, or on arXiv in the meantime), targets that gap.
Introducing SynthID Bio
As a part of our broader vision for bioresilience, SynthID Bio’s watermarking approach serves as an important, tangible verification layer embedded in the biological design itself.

A big-tent or small-tent AI safety movement?
In recent weeks, two narratives about AI safety have emerged: either AI existential risk is real and imminent, or AI leaders’ and whistleblowers’ claims to that effect are insincere — a “psyop” or hype or a twisted form of regulatory capture.
unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
Run and serve Decision Models like Laya (open-source Jev) locally

We tested our own WAF with frontier AI models. Here’s what we found
“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.

Anthropic is cutting off its internal evaluations from the internet
After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations.

2026 Usage Policy update
Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers.
Quoting Victoria Kim
Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said.

When the Safety Test Became the Threat: The Machine That Found Its Own Way Out
OpenAI built a room with no doors – or so it thought.
Expanding the Cyber Verification Program
We’re launching a new, expanded version of our Cyber Verification Program (CVP), which makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals.
Sophos cuts threat investigation time by 96% with OpenAI Daybreak
Discover how Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96% and automate 52% of MDR cases while preserving human oversight.

Launching an opt-in vulnerability-finding service for open-source software
We’re launching OSS Scanner, an opt-in vulnerability scanner for the open-source ecosystem informed by our experience using Claude to find vulnerabilities during Project Glasswing.

Ukraine’s drones knock out AI data center belonging to "Russia’s Google"
Ukrainian drone strikes have knocked out two of five data centers belonging to the Russian tech giant Yandex.
Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.
Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect
Three fired OpenAI safety researchers dispute allegations of mishandling sensitive information, warning in an open letter that their dismissals are creating a chilling effect on the company’s AI safety culture.

Amazon and others are done keeping data center deals secret. Is it enough to build trust?
Amazon says it will stop using NDAs when negotiating data center deals with local governments, following a similar move from Microsoft earlier this year.
Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user.
How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper.