
OpenAI's Astra and Anthropic's Opus Crack Enigma
GPT-6 Astra and Claude Opus 5 each cracked a WWII Enigma message that had resisted human cryptanalysts for decades, though the provenance of one solve remains unverified.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

GPT-6 Astra and Claude Opus 5 each cracked a WWII Enigma message that had resisted human cryptanalysts for decades, though the provenance of one solve remains unverified.

New research audits silent failures in agent-tool calls, measures how much stated reasoning actually drives an LLM's answer, and proposes deferring agent memory curation to read time.

Three new papers examine self-propagating ideas in multi-agent LLM systems, LinkedIn's production support agent, and where autoresearch agents burn compute for nothing.

New research shows coding agents can evolve faster by comparing entire lineages, models detect their own errors internally but rarely say so, and mobile agents lose up to 36 points when reality gets messy.

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

Three new arXiv papers show alignment faking needs no instrumental incentive, LLM scheming spikes in low-resource languages, and Microsoft researchers map the psychological risks of everyday chatbot use.

Three new papers show LLM answers flip under paraphrasing, coding-agent harnesses skew benchmarks more than models do, and a cloud provider ran agents safely for eight months with layered access control.

This week's research roundup covers agent benchmarks that reward exploits over real capability, reasoning models that give up despite having the answer, and why LoRA can't internalize multi-step procedures.

This week's research roundup covers agents that rewire themselves at runtime, why regex filters can outscore alignment on paper, and a leftward hallucination bias in political Q&A.

Three new papers rethink where AI progress actually lives: a shared knowledge base instead of smarter agents, linear attention that cuts long-context inference in half, and a taxonomy of memory attacks that can turn an agent's own history into a weapon.

AlayaWorld is a 15B open-weight video diffusion world model from Alaya Lab that sustains interactive, camera-controllable environments past 60 seconds.

Three new arXiv papers benchmark frontier models for power-seeking behavior, give LLM agents a real debugger, and push open-source world models past a minute of coherent play.