
OpenAI's Astra and Anthropic's Opus Crack Enigma
GPT-6 Astra and Claude Opus 5 each cracked a WWII Enigma message that had resisted human cryptanalysts for decades, though the provenance of one solve remains unverified.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

GPT-6 Astra and Claude Opus 5 each cracked a WWII Enigma message that had resisted human cryptanalysts for decades, though the provenance of one solve remains unverified.

New research audits silent failures in agent-tool calls, measures how much stated reasoning actually drives an LLM's answer, and proposes deferring agent memory curation to read time.

OpenAI's original cost-efficient GPT-5 variant pairs a 400K context window with $0.25/$2.00 per million token pricing, still doing quiet duty as a cheap backbone for research agents a year after launch.

Three new papers show LLM answers flip under paraphrasing, coding-agent harnesses skew benchmarks more than models do, and a cloud provider ran agents safely for eight months with layered access control.

This week's research roundup covers agent benchmarks that reward exploits over real capability, reasoning models that give up despite having the answer, and why LoRA can't internalize multi-step procedures.

Anthropic's July 2026 release prices near-Fable-5 coding and agentic performance at Opus 4.8 rates, doubling Frontier-Bench scores and landing within 0.5 points of Fable 5 on CursorBench at half the cost.

DeepSeek-R1 is the 671B-parameter open-weight reasoning model that matched OpenAI o1 on math and coding benchmarks and triggered a $589 billion single-day drop in Nvidia's market cap in January 2025.

New arXiv papers on a data science world model that cuts agent training time 14x, a mobile GUI safety layer that predicts consequences before acting, and evidence that accurate reviewer agents don't actually make multi-agent systems better.

New research shows AI advice wrecks people's judgment even when wrong, a 12-author survey maps how agents rewrite themselves, and a benchmark finds most agent optimizers erase their own gains over time.

Snowflake's reasoning-first text-to-SQL model tops the BIRD benchmark at 71.83% execution accuracy, trained with GRPO and a reward that only checks if the SQL runs correctly.

Google DeepMind's Gemini 3 Pro debuted at 1501 Elo on LMArena with 91.9% on GPQA Diamond and a 1M-token context window, before Google retired it for Gemini 3.1 Pro.

Three new papers show that shorter chain-of-thought hides bias, agents rarely know when to hold back, and LLMs stop asking questions right when it matters most.