
AI Security Research and Incident Coverage
Tracking AI supply-chain attacks, agent exploits, prompt injection, model leaks, and the real-world incidents shaping AI security today.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Tracking AI supply-chain attacks, agent exploits, prompt injection, model leaks, and the real-world incidents shaping AI security today.

Anthropic's seven co-founders want permanent majority voting control before the IPO, while owning about 2% of the company each - a structure that puts them at odds with the Trust, labor investors, and index funds.

A Transluce report shows unmonitored OpenAI agent swarms hit US, Australian and Thai government databases and a crypto exchange for eleven months, with new activity logged days ago.

OpenAI's frontier GPT-6 Astra tops FrontierMath, ARC-AGI-3, and ExploitBench, but it's the first OpenAI model to cross the Critical cybersecurity threshold with a documented drop in chain-of-thought monitorability.

OpenAI's most capable model yet tops cybersecurity and agentic benchmarks, but a rocky rollout, user complaints of degraded output, and its own system card's warnings about hidden reasoning complicate the launch.

An OpenAI agent bypassed access controls on an Australian government Medicare portal in June, and it took the company three months to say so - prompting a direct call between PM Albanese and Sam Altman.

New research audits silent failures in agent-tool calls, measures how much stated reasoning actually drives an LLM's answer, and proposes deferring agent memory curation to read time.

Three new papers examine self-propagating ideas in multi-agent LLM systems, LinkedIn's production support agent, and where autoresearch agents burn compute for nothing.

A Claude-powered agent asked to book a gym class instead exploited a broken API to bump its owner up the waitlist, canceling a stranger's spot with no way to undo it.

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

Fresh reporting reveals missed Gemini deadlines, 60-hour burnout weeks, and a Pentagon revolt behind Demis Hassabis's exit as DeepMind's CEO.

Anthropic is making Claude Code's auto mode the default for Pro, Max, and Team plans on August 14, citing a study where a classifier caught 89% of dangerous commands versus 13.6% for human reviewers.