
Best AI Models for Agentic Tool Use - August 2026
Claude Opus 5 tops independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% for a third of the price.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Claude Opus 5 tops independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% for a third of the price.

Three new papers examine self-propagating ideas in multi-agent LLM systems, LinkedIn's production support agent, and where autoresearch agents burn compute for nothing.

New research shows coding agents can evolve faster by comparing entire lineages, models detect their own errors internally but rarely say so, and mobile agents lose up to 36 points when reality gets messy.

A Claude-powered agent asked to book a gym class instead exploited a broken API to bump its owner up the waitlist, canceling a stranger's spot with no way to undo it.

A step-by-step guide to using ChatGPT, Claude, or an AI agent app to write negotiation scripts and lower your internet, phone, and subscription bills.

Updated August 10: Anthropic's Agent SDK credit plan died before launch, Claude Managed Agents adds a new session-hour billing line, and E2B, Modal, and Daytona all rebuilt their per-second pricing.

Meta's 30B open-weight local agent model beats its closest open rivals on independent tool-use tests, but trails on long agent sessions and on prompt-injection resistance.

Meta released Muse Glimmer, a 30B open-weight model distilled from Muse Spark that runs on a single consumer GPU, reversing its April pivot toward closed frontier models.

Cyera will pay about $1 billion for Oasis Security, its fifth 2026 acquisition, as enterprises scramble to manage the credentials of AI agents outnumbering human employees 45 to 1.

Microsoft says MAI-Cyber-1-Flash helps MDASH beat every rival on the CyberGym benchmark, but the score isn't on CyberGym's own public leaderboard, and Wiz topped it the same day with a lower, verified number.

Three new papers show LLM answers flip under paraphrasing, coding-agent harnesses skew benchmarks more than models do, and a cloud provider ran agents safely for eight months with layered access control.

This week's research roundup covers agent benchmarks that reward exploits over real capability, reasoning models that give up despite having the answer, and why LoRA can't internalize multi-step procedures.