
Best AI Models for Agentic Tool Use - August 2026
Claude Opus 5 tops independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% for a third of the price.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Claude Opus 5 tops independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% for a third of the price.

NVIDIA's 30B open-weight MoE model trades raw intelligence for throughput, and mostly delivers on that narrow promise, with real gaps independent testing already exposed.

NVIDIA distilled its 550B Nemotron 3 Ultra down to a 30B MoE model with 3B active parameters, aimed at the boring, high-volume work inside agent pipelines.

NVIDIA's 30B MoE model with 3B active parameters, distilled from Nemotron 3 Ultra, hits 86% PinchBench accuracy at up to 4x the output speed of comparable open models.

Alibaba's 2.4 trillion parameter flagship ships with real pricing and a published benchmark table, but the open-weight release it promised for this week still hasn't shown up.

Meta's 30B open-weight agentic model distilled from Muse Spark runs on a single consumer GPU, ships under Apache 2.0, and leads Gemma4-31B and Qwen3.6-27B on 5 of 8 agentic benchmarks.

Meta's second Muse model ships a public API at $1.25/$4.25 per million tokens, a 1M-token context window, and the top score on Meta's own tool-use benchmarks.

Meta's second Muse Spark ships with a real API, a 1M-token context window and the cheapest pricing among frontier-class agents, but only US developers can touch it.

Cyera will pay about $1 billion for Oasis Security, its fifth 2026 acquisition, as enterprises scramble to manage the credentials of AI agents outnumbering human employees 45 to 1.

Microsoft's first in-house cybersecurity model is a 137B sparse MoE fine-tune that drives its MDASH vulnerability harness to a self-reported 95.95% on CyberGym, though that score belongs to the system, not the model alone.

Google DeepMind's cheapest paid Gemini tier prices input at $0.30/M and output at $2.50/M tokens, more than doubling OSWorld-Verified and Terminal-Bench 2.1 scores over Gemini 3.1 Flash-Lite while trailing GPT-5.4 mini on raw coding benchmarks.

Claude Opus 5 ties Claude Fable 5 on independent benchmarks at roughly a quarter of the cost, though a rough launch week and cybersecurity limits keep it from being an unqualified win.