Best AI Models for Agentic Tool Use - August 2026

Claude Opus 5 tops independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% for a third of the price.

Agentic Tool Use Top: Claude Opus 5 Updated monthly
Best AI Models for Agentic Tool Use - August 2026

TL;DR

  • Claude Opus 5 tops independent SWE-bench Verified tracking at 96% - though Anthropic's own launch materials skip that benchmark entirely for custom evals
  • Qwen3.8 Max leads OSWorld-Verified computer use at 86.1% while charging $2/$6 per million tokens, a fraction of what Claude or GPT models cost
  • Best value goes to Claude Sonnet 5: 85.2% SWE-bench Verified at $2/$10 introductory pricing, which expires August 31, 2026

Four months ago, this page named Claude Opus 4.6 the agentic leader at 80.8% on SWE-bench Verified. Every model in that table has since been surpassed, several of them twice. The benchmark itself is nearly saturated at the top: the current top three entries sit within one point of each other.

As of August 2026, Claude Opus 5 holds the strongest position on independently tracked SWE-bench Verified scores at 96%, per aggregators llm-stats.com and BenchLM. But Anthropic's own Opus 5 announcement doesn't cite that number at all - it leans on Frontier-Bench, CursorBench, and ARC-AGI-3 instead. On computer use specifically, Qwen3.8 Max leads OSWorld-Verified at 86.1%, ahead of every Claude and GPT model tracked, at a quarter of Opus 5's price.

That split matters more than any single leaderboard position. Labs are increasingly reporting scores on benchmarks they designed themselves, which makes cross-model comparison harder than it was even three months ago.

Rankings Table

RankModelProviderSWE-bench VerifiedOSWorld-VerifiedPrice (In/Out per 1M)Verdict
1Claude Opus 5Anthropic96%*-$5/$25Best combined agentic leader on independent trackers
2Claude Mythos 5Anthropic95.5%85%$10/$50Restricted to Glasswing partners, not a practical pick
3Claude Fable 5Anthropic95%85%$10/$50Most capable public Claude, globally available since July 1
4Claude Opus 4.8Anthropic88.6%83.4%$5/$25Prior-gen flagship, still podium OSWorld months later
5GPT-5.6 SolOpenAI-83.2%†$5/$30Terminal-Bench and BrowseComp leader
6Qwen3.8 MaxAlibaba-86.1%$2/$6OSWorld-Verified #1 at a fraction of Claude's price
7GPT-5.5OpenAI-78.7%$5/$30Balanced all-rounder, first fully retrained GPT-5 base
8Claude Sonnet 5Anthropic85.2%81.2%$3/$15 (intro $2/$10 thru Aug 31)Best value near-Opus agentic performance
9Gemini 3.5 FlashGoogle DeepMind-78.4%$1.50/$9MCP Atlas leader among mainstream models, fastest output
10DeepSeek V4DeepSeek80.6%-$1.74/$3.48Strongest open-weight SWE-bench score, MIT license
11Kimi K3Moonshot AI--$3/$15Largest open-weight model shipped, strong MCP and terminal scores
12MiniMax M3MiniMax--$0.60/$2.40Cheapest capable option, benchmarks are self-reported

Dashes show no independently verified score under that exact benchmark version. *Opus 5's SWE-bench Verified score comes from third-party trackers, not from Anthropic's own release materials. †GPT-5.6 Sol's OSWorld-Verified figure is from Alibaba's own comparison chart at Qwen3.8 Max's launch, not yet confirmed on an independent OSWorld tracker. Our agentic benchmark leaderboard tracks additional models as new data publishes.

Alibaba's Qwen3.8-Max benchmark comparison table showing scores across FrontierSWE, PaperBench, and general agent benchmarks Alibaba's own launch comparison table for Qwen3.8-Max, spanning coding and general-agent benchmarks against Claude and GPT models. Self-reported scores like these still need independent confirmation. Source: marktechpost.com

Detailed Analysis

Claude Opus 5 - Leading a Benchmark It Doesn't Cite

Claude Opus 5 launched July 24, 2026, priced identically to its predecessor Opus 4.8 at $5 input / $25 output per million tokens. Anthropic's pitch was unusually modest for a flagship: the company still calls Claude Fable 5 its most capable model, positioning Opus 5 as a way to get most of that capability at a lower cost per task.

The benchmarks Anthropic actually disclosed tell an odd story next to the SWE-bench Verified number that made this page's headline. Frontier-Bench v0.1 more than doubled versus Opus 4.8 (43.3% against 18.7%), and CursorBench 3.2 landed within half a point of Fable 5's ceiling (70.0% versus 70.5%) at roughly half the cost per attempt. Neither of those is SWE-bench Verified. The 96% figure driving Opus 5's #1 spot on this table comes completely from independent trackers running their own harness against the model, not from anything Anthropic published at launch.

That's worth sitting with. When a lab stops citing the field's standard benchmark and leads with proprietary ones instead, it's either because the standard benchmark is saturated and no longer differentiates, or because the number doesn't flatter the release as much as the custom ones do. Both explanations are plausible here given how tightly the top three SWE-bench Verified scores now cluster.

When a lab drops the standard benchmark for custom ones at launch, that's either saturation or an unflattering number. Both are live possibilities for Opus 5.

Read the full Claude Opus 5 review for hands-on notes on the new "xhigh" reasoning effort tier and 1M-token context window.

Qwen3.8 Max - Computer Use at a Quarter of the Price

Alibaba shipped Qwen3.8 Max to general availability on August 3, 2026, and its OSWorld-Verified score of 86.1% is currently the highest confirmed on that leaderboard - ahead of Claude Mythos 5 and Fable 5 at 85%, and Claude Opus 4.8 at 83.4%. At $2 input / $6 output per million tokens, it costs a fraction of any Claude model in this table.

The catch sits outside the benchmark table. Alibaba promised open weights for "the week of August 10." As of this article's publication - two days past that window - the weights still hadn't appeared on Hugging Face or ModelScope. Qwen3.8 Max also trails badly on coding specifically: 67.7% on SWE-bench Pro against Fable 5's 80.0%, a 12-point gap that OSWorld leadership doesn't close. Pick it for computer-use and browser-automation agents; look elsewhere for repository-scale coding work.

GPT-5.6 Sol - Terminal and Browsing Specialist

GPT-5.6 Sol reached general availability July 9, 2026, after a two-week preview restricted to roughly 20 government-vetted partners - a first for a frontier model launch. Its strengths are narrower and sharper than Opus 5's or Qwen3.8 Max's: 88.8% on Terminal-Bench 2.1 (91.9% in Ultra mode, which coordinates multiple Sol subagents on split tasks) and 92.2% on BrowseComp, both field-leading.

OpenAI's own comparison for GPT-5.6's smaller sibling puts its OSWorld 2.0 score at 62.6% - a newer, harder OSWorld revision not directly comparable to the OSWorld-Verified scores elsewhere in this table. Sol reportedly hits that number using 85% fewer output tokens than Claude Opus 4.8 needs for equivalent tasks, which matters more for production cost than the headline percentage does.

Claude Sonnet 5 - The Value Pick, For Three More Weeks

Claude Sonnet 5 is the practical recommendation for most teams building agents today. It scores 85.2% on SWE-bench Verified and 81.2% on OSWorld-Verified - within a few points of the Opus tier - while its introductory pricing of $2 input / $10 output per million tokens undercuts standard Sonnet pricing by a third. That discount expires August 31, 2026, reverting to $3/$15.

On BrowseComp, the agentic-search benchmark that measures whether a model can dig up hard-to-find information autonomously, Sonnet 5 scores 84.7% single-agent - trailing only GPT-5.5's 84.4% among broadly available models and well clear of its own predecessor, Claude Sonnet 4.6, at 76.2%.

The Open-Weight Tier Closes the Gap

Open-weight models no longer trail the field by the 20-point margin this page reported in April. DeepSeek V4 Pro posts 80.6% on SWE-bench Verified at $1.74/$3.48 per million tokens under a MIT license - within 8 points of Claude Opus 5's independently tracked score, at roughly a third of the price. Kimi K3, Moonshot AI's 2.8 trillion-parameter release, shipped its weights on Hugging Face July 26, making it the largest open-weight model released to date; it posts strong showings on MCP Atlas (84.2%) and Terminal-Bench 2.1 (88.3%), though it lacks a confirmed SWE-bench Verified figure from an independent tracker.

MiniMax M3 undercuts everyone on price at $0.60/$2.40 per million tokens with weights already live on Hugging Face, but its 59.0% SWE-bench Pro score is still self-reported and hasn't been independently reproduced. Treat it as a budget option worth testing against your own workload, not a settled ranking.

Methodology

Four benchmarks anchor this ranking, each measuring a different slice of agentic capability:

SWE-bench Verified tests 500 real GitHub issues across a dozen Python repositories, requiring end-to-end patch generation that passes held-out tests. It's the closest thing the field has to a standard, though the top scores are now clustered within a single point of each other - a sign the benchmark is approaching saturation for frontier models.

OSWorld-Verified scores autonomous computer use across 369 real desktop and web tasks on Ubuntu, Windows, and macOS - file management, multi-app workflows, and terminal operations graded against actual machine state rather than self-reported output. See the computer use leaderboard for the full breakdown.

Terminal-Bench 2.1 measures command-line agent workflows requiring planning, iteration, and tool coordination across a real shell. Our terminal benchmark leaderboard tracks additional entries.

MCP Atlas assesses tool-use competency through the Model Context Protocol specifically: 36 real MCP servers, 220 tools, and roughly 1,000 multi-step tasks scored against atomic factual claims grounded in tool outputs. See the MCP server ecosystem leaderboard.

Function calling in the narrow sense - parameter extraction and API call formatting - has become a separate story. On the Berkeley Function Calling Leaderboard v4, Qwen3.7 Max currently leads at 75.0%, with Qwen3.7 Plus close behind at 72.9%, according to a July 2026 BenchLM snapshot. Neither model tops this page's main table, reinforcing that narrow function calling and end-to-end agentic task completion are different skills that don't move together.

Tau2-bench, which tests tool-agent-user interaction in realistic customer service dialogues, shows near-saturation in the telecom domain: GLM-5.2 and a smaller specialist model both clear 99% on Artificial Analysis's tracker. That domain is effectively solved; the retail and voice domains added to the 2026 benchmark revision are not.

The caveat from prior updates to this page still holds: scaffold and harness choice can shift agentic scores by double digits, while swapping the underlying model moves scores by a fraction of that. The numbers above describe model capability under each benchmark's standardized conditions - production performance depends on the tool loop, retry logic, and context management wrapped around whichever model you choose.

Historical Progression

  • February 2026 - Claude Opus 4.6 reached 72.7% on OSWorld and 80.8% on SWE-bench Verified, the top scores this page tracked at the time.

  • April 2026 - DeepSeek V4 and GPT-5.5 launched the same day, April 24. MiniMax M2.5 briefly cracked the SWE-bench Verified top-5, the first non-Anthropic, non-Google, non-OpenAI model to do so.

  • May 2026 - Gemini 3.5 Flash launched at Google I/O; Claude Opus 4.8 followed on May 28 with 88.6% SWE-bench Verified and 83.4% OSWorld-Verified.

  • June 2026 - Claude Fable 5 and Mythos 5 launched June 9 as Anthropic's first Mythos-class models, then were pulled offline globally by an US export control order on June 12. Claude Sonnet 5 shipped June 30.

  • July 2026 - Export controls on Fable 5 and Mythos 5 lifted July 1; Fable 5 returned globally while Mythos 5 stayed limited to Glasswing partners. GPT-5.6 Sol reached general availability July 9. Kimi K3 launched July 16 and shipped open weights July 26. Claude Opus 5 launched July 24.

  • August 2026 - Qwen3.8 Max reached general availability August 3, taking the OSWorld-Verified lead at 86.1% - still without the open weights Alibaba promised for launch week.

Six months ago, OSWorld's ceiling sat in the low 70s and SWE-bench Verified topped out around 81%. Both benchmarks now show frontier models clearing 85-96%, with the gap to open-weight alternatives closing from over 20 points to single digits on at least one major eval.


FAQ

What's the best model for agentic tool use right now?

Claude Opus 5 leads independent SWE-bench Verified tracking at 96%, but Qwen3.8 Max tops OSWorld-Verified computer use at 86.1% for a quarter of the price. The right pick depends on whether your agent codes or operates a desktop.

What's the cheapest model that's still competitive?

MiniMax M3 at $0.60/$2.40 per million tokens is the cheapest option with open weights already live, though its benchmark scores remain self-reported. DeepSeek V4 at $1.74/$3.48 has independently confirmed scores.

Is open-source competitive for agentic tool use now?

Much more than it was in April. DeepSeek V4 sits within 8 points of Claude Opus 5's SWE-bench Verified score at roughly a third of the price, and Kimi K3 is the largest open-weight model ever shipped, with weights live since July 26.

Why doesn't Claude Opus 5's SWE-bench score come from Anthropic directly?

Anthropic's launch materials cite Frontier-Bench, CursorBench, and ARC-AGI-3 instead of SWE-bench Verified. The 96% figure on this page comes from independent trackers running their own harness against the model, not from Anthropic's release.

Which model is best for computer use specifically?

Qwen3.8 Max leads OSWorld-Verified at 86.1%, ahead of Claude Mythos 5 and Claude Fable 5 at 85% each. Its promised open weights, however, are still unreleased past Alibaba's own deadline.

How often do agentic rankings change?

Fast - multiple reshuffles happened in the four months before this update. MiniMax M2.5 appeared in the SWE-bench top-5 with no warning. Check lastVerified at the top of this page and the agentic benchmarks leaderboard directly.


Sources:

✓ Last verified August 12, 2026

LinkedIn
Reddit
Hacker News
Telegram
Best AI Models for Agentic Tool Use - August 2026
About the author AI Benchmarks & Tools Analyst

James is a software engineer turned tech writer who spent six years building backend systems at a fintech startup in Chicago before pivoting to full-time analysis of AI tools and infrastructure.