
GPT-5 mini
OpenAI's original cost-efficient GPT-5 variant pairs a 400K context window with $0.25/$2.00 per million token pricing, still doing quiet duty as a cheap backbone for research agents a year after launch.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

OpenAI's original cost-efficient GPT-5 variant pairs a 400K context window with $0.25/$2.00 per million token pricing, still doing quiet duty as a cheap backbone for research agents a year after launch.

Three new papers show LLM answers flip under paraphrasing, coding-agent harnesses skew benchmarks more than models do, and a cloud provider ran agents safely for eight months with layered access control.

This week's research roundup covers agent benchmarks that reward exploits over real capability, reasoning models that give up despite having the answer, and why LoRA can't internalize multi-step procedures.

Anthropic's July 2026 release prices near-Fable-5 coding and agentic performance at Opus 4.8 rates, doubling Frontier-Bench scores and landing within 0.5 points of Fable 5 on CursorBench at half the cost.

DeepSeek-R1 is the 671B-parameter open-weight reasoning model that matched OpenAI o1 on math and coding benchmarks and triggered a $589 billion single-day drop in Nvidia's market cap in January 2025.

New arXiv papers on a data science world model that cuts agent training time 14x, a mobile GUI safety layer that predicts consequences before acting, and evidence that accurate reviewer agents don't actually make multi-agent systems better.

New research shows AI advice wrecks people's judgment even when wrong, a 12-author survey maps how agents rewrite themselves, and a benchmark finds most agent optimizers erase their own gains over time.

Snowflake's reasoning-first text-to-SQL model tops the BIRD benchmark at 71.83% execution accuracy, trained with GRPO and a reward that only checks if the SQL runs correctly.

Google DeepMind's Gemini 3 Pro debuted at 1501 Elo on LMArena with 91.9% on GPQA Diamond and a 1M-token context window, before Google retired it for Gemini 3.1 Pro.

Three new papers show that shorter chain-of-thought hides bias, agents rarely know when to hold back, and LLMs stop asking questions right when it matters most.

Nous Research's 36B open-weight model matches Hermes 4 70B on most benchmarks, tops RefusalBench on alignment, and is the first production model trained entirely on the Solana-secured Psyche network.

Google's hybrid reasoning workhorse pairs a 1M-token context window with $0.30/$2.50 per million token pricing and a toggleable 0-24,576 token thinking budget, now heading toward an October 2026 shutdown.