
Best AI Models for Agentic Tool Use - August 2026
Claude Opus 5 tops independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% for a third of the price.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Claude Opus 5 tops independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% for a third of the price.

Meta's second Muse model ships a public API at $1.25/$4.25 per million tokens, a 1M-token context window, and the top score on Meta's own tool-use benchmarks.

A single chat message could break an AI agent out of Claude Cowork's isolated VM and reach an entire Mac, and Anthropic closed the report as informative.

New arXiv research measures context quality as a leading indicator of agent reliability, gives computer-use agents a more reliable execution layer, and catches coding agents that covertly sabotage their own guardrails.

A step-by-step guide to setting up your first AI browser agent, giving it a real task, and using it safely without handing over your passwords.

H Company's open-weight sparse MoE vision-language model purpose-built for desktop computer use, scoring 82.6% on OSWorld-Verified with only 3B active parameters.

Claude Fable 5 leads OSWorld-Verified at 85% after its 19-day US suspension ended July 1 - Holo3 open-source at 82.6% and Claude Sonnet 5 at $2/M tokens reshape the value calculus.

Anthropic's Sonnet 5 is the first mid-tier model that genuinely competes with Opus-class agents on coding and computer use, released June 30 at $2/$10 per million tokens.

Anthropic's latest Sonnet-class model brings near-Opus coding performance to mid-tier pricing, with major agentic search and computer use gains over Sonnet 4.6.

Sakana Fugu tops SWE-Bench Pro by routing tasks across rival LLMs, Microsoft's 9B browser agent beats OpenAI Operator, and a 3B model from Weibo matches DeepSeek V3.2 on math.

Microsoft Research's family of open-weight browser computer use agents (4B, 9B, 27B) that beat OpenAI Operator and Gemini 2.5 Computer Use on Online-Mind2Web.

Three new papers: agents that compile runs into 8-13x faster state machines, benchmark scores that shift with compute budget, and big brands monopolizing LLM recommendations.