Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

TechCrunch: AI4 sources2d ago

Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide

An Anthropic AI model provided false information about an unsolved homicide to a Philadelphia Police Department (PPD) tipline, according to a report from 6abc.

2 outlets
The Verge: AI4 sources13h ago

Satya Nadella says we should assume all AI models are ‘compromised’

In a lengthy post on X, Microsoft's CEO laid out his views on the dangers posed by highly advanced AI models and how to confront those risks.

3 outlets
Simon Willison's Weblog2 sources4d ago

OpenAI “rogue” agent activities found on Wikimedia projects

OpenAI “rogue” agent activities found on Wikimedia projects

2 outlets
Ars Technica: AI2 sources3d ago

Google rolls out improved SynthID AI content detector, now available globally

Every piece of AI content from Google's Gemini models has a hidden SynthID label, which makes it very difficult to pass the content off as authentic.

2 outlets
OpenAI News2 sources5d ago

Our approach to EU text provenance rules

How OpenAI is approaching text watermarking under EU rules.

Anthropic Research2d ago

Investigating unintended model actions in our evaluations and internal use

Investigating unintended model actions in our evaluations and internal use

Anthropic News3d ago

Introducing the Anthropic Cyber Mission

Today we’re launching the Anthropic Cyber Mission, a long-term commitment to securing the systems everyone depends on.

Vendor claim only
MarkTechPost15h ago

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context

OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research.

GitHub: pydantic/pydantic-ai2d ago

pydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)

<!-- Release notes generated using configuration in .github/release.yml at main -->

Code
WIRED: AI10h ago

I Made Terrible Games With Google’s AI Playground

A long day’s haul in the video game slop mines.

Latent Space2d ago

[AINews] not much happened today

More AI safety intrigue in the alignment below.

The Decoder1d ago

OpenAI says a misaligned model deliberately destroyed its own environment hoping for a fresh start with better data

OpenAI has documented new cases of misaligned model behavior.

Video
AI Engineer (YouTube)8h ago

AI Security Engineer Foundations + Certificate — Micah Silverman, Snyk

Micah Silverman uses that capture the flag exercise to connect AI risks with familiar application security controls.

Paper
Hugging Face Daily Papers2 sources3d ago

Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents

In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope.

Paper
GitHub: BerriAI/litellm12h ago

BerriAI/litellm v1.105.0

Verify using the pinned commit hash (recommended):

Code
Epoch AI: Gradient Updates2d ago

Can AI automate AI R&D yet?

Existing evidence shows that AI can perform software engineering tasks relevant to AI research, dataset creation, open-ended optimization of defined metrics (autoresearch), and more.

Google Research Blog9d ago

Toward provably private learning from federated data

In 2017, Google introduced Federated Learning (FL) a machine learning technique that trains models across decentralized, private data.

Vendor claim only
GitHub: anthropics/claude-code3d ago

anthropics/claude-code v2.1.294

Fixed prompt and agent hooks written as instructions (such as "Block commands that...") allowing what they should block

Code
Microsoft Research Blog5d ago

What AI gets wrong and what failure teaches us

Jennifer Neville is a partner research manager at Microsoft who’s built a career around understanding and advancing AI for real-world use, and much like the human-AI interactions she’s been studying, her early-career path was multiturn: math, then physics; cognitive science, then work; and finally computer science—despite her best efforts to avoid the field.

Last Week in AI11d ago

Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots

Anthropic and OpenAI race to release smarter and cheaper models

NVIDIA Technical Blog4d ago

Validate AI Factory Changes with Digital Twins and AI Agents

AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration...AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration services, security controls, and a rapidly changing software stack.

GitHub: huggingface/trl5d ago

huggingface/trl v1.14.2

Patch release fixing two cases of silently wrong training and three crashes.

Code
AWS Machine Learning Blog5d ago

Responsible AI governance: How AWS positions customers to align with ISO/IEC 42005:2025

With generative AI adoption moving faster than the personal computer or the internet and global AI-related investment in 2025 representing $581.69 billion, organizations must position their workforce to use AI to power their operations while employing it responsibly.

IEEE Spectrum: AI6d ago

Attempts to Keep Humans in the AI Loop May Actually Push Them Out

A crucial safeguard against AI agents going rogue—keeping humans in the loop to review and approve their decisions—will fail unless designers and users change their current practices, a trio of leading AI ethics researchers argue.

GitHub: langchain-ai/langgraph5d ago

langchain-ai/langgraph sdk==0.4.6: langgraph-sdk==0.4.6

fix(sdk-py): percent-encode threadid and assistantid in thread stream requests (#9213)

Code
GitHub Blog: AI and ML4d ago

Secret protection must scale with software

Today, one in three pull requests on GitHub involves an AI agent.

Interconnects (Nathan Lambert)5d ago

The Cyber Risk Discourse is Broken

Western voices saying open weight models are necessary for defense and banning them will make the world less safe, occupied by AI risk moderates to different extremes.

Video
AI Explained (YouTube)10d ago

OpenAI Security: Controlling Models is Now ‘Hell’

A cracked cipher, an OpenAI security warning, Gemini 4 Argon, RSI paper (co-authored by a who’s who of AI), Lab White House commitments, new hacks emerging, ‘deep personas’, biology Kasparov competitions, and so much more, ending with an epic Opus outro.

MIT Technology Review (AI)6d ago

People really hate AI, so why can’t they get enough?

Over the summer I talked to the CEO of Springboards, a startup building an LLM that’s designed to come up with a wider variety of responses than its mainstream rivals do.

Hugging Face Blog12d ago

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

Our latest paper, ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (read it on Hugging Face, or on arXiv in the meantime), targets that gap.

Google DeepMind Blog11d ago

Introducing SynthID Bio

As a part of our broader vision for bioresilience, SynthID Bio’s watermarking approach serves as an important, tangible verification layer embedded in the biological design itself.

Vendor claim only
AI Snake Oil10d ago

A big-tent or small-tent AI safety movement?

In recent weeks, two narratives about AI safety have emerged: either AI existential risk is real and imminent, or AI leaders’ and whistleblowers’ claims to that effect are insincere — a “psyop” or hype or a twisted form of regulatory capture.

GitHub: unslothai/unsloth13d ago

unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library

Run and serve Decision Models like Laya (open-source Jev) locally

Code
Cloudflare Blog: AI12d ago

We tested our own WAF with frontier AI models. Here’s what we found

“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.

TechCrunch: AI2 sources1d ago

Anthropic is cutting off its internal evaluations from the internet

After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations.

2 outlets
Anthropic News3d ago

2026 Usage Policy update

Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers.

Simon Willison's Weblog4d ago

Quoting Victoria Kim

Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said.

MarkTechPost1d ago

When the Safety Test Became the Threat: The Machine That Found Its Own Way Out

OpenAI built a room with no doors – or so it thought.

Anthropic News5d ago

Expanding the Cyber Verification Program

We’re launching a new, expanded version of our Cyber Verification Program (CVP), which makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals.

Vendor claim only
OpenAI News2d ago

Sophos cuts threat investigation time by 96% with OpenAI Daybreak

Discover how Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96% and automate 52% of MDR cases while preserving human oversight.

Anthropic Research3d ago

Launching an opt-in vulnerability-finding service for open-source software

We’re launching OSS Scanner, an opt-in vulnerability scanner for the open-source ecosystem informed by our experience using Claude to find vulnerabilities during Project Glasswing.

Vendor claim only
Ars Technica: AI2d ago

Ukraine’s drones knock out AI data center belonging to "Russia’s Google"

Ukrainian drone strikes have knocked out two of five data centers belonging to the Russian tech giant Yandex.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes

This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.

Paper
TechCrunch: AI3d ago

Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect

Three fired OpenAI safety researchers dispute allegations of mishandling sensitive information, warning in an open letter that their dismissals are creating a chilling effect on the company’s AI safety culture.

TechCrunch: AI2d ago

Amazon and others are done keeping data center deals secret. Is it enough to build trust?

Amazon says it will&#160;stop using NDAs&#160;when negotiating data center deals with local governments, following&#160;a&#160;similar move from Microsoft&#160;earlier this year.

Paper
Hugging Face Daily Papers3d ago

Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub

Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense

Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper.

Paper