
AI Security Research and Incident Coverage
Tracking AI supply-chain attacks, agent exploits, prompt injection, model leaks, and the real-world incidents shaping AI security today.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Tracking AI supply-chain attacks, agent exploits, prompt injection, model leaks, and the real-world incidents shaping AI security today.

Apple filed suit against OpenAI on July 10, alleging a coordinated scheme to steal iPhone trade secrets through recruiting tactics and unauthorized file downloads.

AI agents reproduce 72% of human research ideological bias, lie detectors improve with model scale, and Mastermind beats iterative vulnerability agents by 7 points.

Anthropic accused Alibaba's Qwen lab of running 25,000 fraudulent accounts that extracted 28.8 million Claude interactions in the largest AI distillation attack on record.

Alibaba classified Claude Code as high-risk software after a researcher alleged version 2.1.91 silently identified Chinese corporate users, escalating a bitter distillation war between the two companies.

Meta has restricted engineers from using Claude Code and Codex, citing training data distillation risk. The policy change exposes a structural mismatch between how AI coding tools work and what enterprise AI labs can tolerate.

OpenAI's GPT-5.5-Cyber found CVE-2026-8390 in Firefox's WebAssembly engine before Pwn2Own Berlin - five of six registered exploit entries withdrew.

OpenAI's GPT-5.5-Cyber is a cybersecurity-specialized fine-tune of GPT-5.5, restricted to vetted defenders through the Daybreak Cyber Partner Program and rated 85.6% on the CyberGym benchmark.

The White House won't lift its ban on Anthropic's Fable 5 until the model can be made jailbreak-proof. Security experts explain why that condition is technically impossible.

A Chinese cybercrime network sold $88/week phishing kits that used Google's own Gemini AI to generate fake sites impersonating banks, carriers, and government agencies at scale.

New research reveals MCP error messages triple agent attack success rates, ranks eight models on sycophancy with Claude scoring best, and finds self-evolving agents make 30-42% false edits.

OpenAI's new Lockdown Mode cuts the network exits that prompt injection attacks use to steal data from ChatGPT - but won't stop malicious instructions from entering the model in the first place.