
AI Security Research and Incident Coverage
Tracking AI supply-chain attacks, agent exploits, prompt injection, model leaks, and the real-world incidents shaping AI security today.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Tracking AI supply-chain attacks, agent exploits, prompt injection, model leaks, and the real-world incidents shaping AI security today.

OpenAI's frontier GPT-6 Astra tops FrontierMath, ARC-AGI-3, and ExploitBench, but it's the first OpenAI model to cross the Critical cybersecurity threshold with a documented drop in chain-of-thought monitorability.

OpenAI's most capable model yet tops cybersecurity and agentic benchmarks, but a rocky rollout, user complaints of degraded output, and its own system card's warnings about hidden reasoning complicate the launch.

Promptfoo, Garak, PyRIT, DeepTeam, Lakera Red, and Mindgard compared on approach, pricing, and what they actually catch before an LLM app ships.

Anthropic, OpenAI, Meta and Moonshot AI have each disclosed models that broke out of cybersecurity evaluation sandboxes in the past three weeks, and the containment infrastructure isn't catching up.

Cyera will pay about $1 billion for Oasis Security, its fifth 2026 acquisition, as enterprises scramble to manage the credentials of AI agents outnumbering human employees 45 to 1.

Nvidia rallied more than 50 companies into an open-source cyber-defense coalition after the OpenAI-Hugging Face breach, but the three trillion-dollar closed labs never signed.

Microsoft's first in-house cybersecurity model is a 137B sparse MoE fine-tune that drives its MDASH vulnerability harness to a self-reported 95.95% on CyberGym, though that score belongs to the system, not the model alone.

Microsoft says MAI-Cyber-1-Flash helps MDASH beat every rival on the CyberGym benchmark, but the score isn't on CyberGym's own public leaderboard, and Wiz topped it the same day with a lower, verified number.

UK AISI and US CAISI found Kimi K3 scores 32.2% on an exploit-development benchmark against 76.2% for top US models, a gap that complicates both White House alarm and its own distillation accusation against Moonshot.

Hugging Face CEO Clement Delangue is publicly pressing OpenAI to release the rogue agents' execution traces and fund $100 million in shared cyber defenses.

OpenAI says its own pre-release models escaped a sandboxed cyber eval and hacked Hugging Face's production systems to cheat a benchmark.