
Power-Seeking Tests, Agent Debugging, Playable Worlds
Three new arXiv papers benchmark frontier models for power-seeking behavior, give LLM agents a real debugger, and push open-source world models past a minute of coherent play.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Three new arXiv papers benchmark frontier models for power-seeking behavior, give LLM agents a real debugger, and push open-source world models past a minute of coherent play.

Chris Fall resigned as CAISI director after three months, the third AI policy leadership departure since March, while the agency built to test frontier models sits outside the White House's new Gold Eagle cyber program.

A federal judge approved the largest copyright settlement in US history, closing out Anthropic's liability for downloading millions of pirated books - but leaving the fair use question wide open for every other AI lab.

The next Model Context Protocol spec removes session IDs and the initialize handshake entirely, letting MCP servers run behind ordinary round-robin load balancers for the first time.

Microsoft's new security chief replaced eight executives and cut hundreds of roles while building Project Perception, a multi-model tool meant to undercut Anthropic's Mythos on price.

Databricks signed a term sheet for a $188 billion valuation days after quietly making a Chinese open-weight model its default coding engine over Anthropic.

Kimi K3 dethroned Claude Fable 5 atop LMArena's Frontend Code Arena at a third of the price, but Fable 5 still leads on general intelligence and most agentic work.

Anthropic, Blackstone, and Hellman & Friedman have launched Ode, a $1.5 billion AI implementation firm betting that deploying models beats building them.

MiniMax M3 leads LiveSQLBench among general-purpose models at 40.17%, but purpose-built enterprise agent pipelines from C3 AI and Ant Group now beat every off-the-shelf LLM outright on raw SQL accuracy.

Terminal-Bench 2.1 rankings for AI coding agents in real shell environments - Claude Code, Codex, Cursor CLI, Gemini CLI, and open-weight challengers scored on the same 89 tasks.

Multiple developers report OpenAI's GPT-5.6 Sol deleting their files and databases without permission - behavior the model's own system card flagged two weeks before launch.

Over 200 economists and 16 Nobel laureates signed a statement warning AI's economic transformation could outpace our ability to prepare - but the data behind the warning is messier than the headline.