James Kowalski

James Kowalski

AI Benchmarks & Tools Analyst

James is a software engineer turned tech writer who spent six years building backend systems at a fintech startup in Chicago before pivoting to full-time analysis of AI tools and infrastructure. His engineering background means he doesn't just read the spec sheet - he runs the benchmarks, profiles the latency, and checks whether the marketing claims hold up under real workloads.

He studied Computer Science at the University of Illinois at Urbana-Champaign, where he first got hooked on natural language processing during a senior research project on sentiment analysis. He later completed a certificate in data journalism from Northwestern's Medill School.

At Awesome Agents, James owns the leaderboards and tool comparison coverage. He maintains the site's benchmark tracking methodology and is the person who actually runs the numbers before publishing any ranking. He is also an open-source advocate and contributes to several projects in the LLM inference space.

Based in Chicago, IL.

Articles by James Kowalski
MAI-Cyber-1-Flash

MAI-Cyber-1-Flash

Microsoft's first in-house cybersecurity model is a 137B sparse MoE fine-tune that drives its MDASH vulnerability harness to a self-reported 95.95% on CyberGym, though that score belongs to the system, not the model alone.

Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite

Google DeepMind's cheapest paid Gemini tier prices input at $0.30/M and output at $2.50/M tokens, more than doubling OSWorld-Verified and Terminal-Bench 2.1 scores over Gemini 3.1 Flash-Lite while trailing GPT-5.4 mini on raw coding benchmarks.

LLM API Pricing Comparison - July 2026

LLM API Pricing Comparison - July 2026

Verified July 27: Claude Opus 5 lands at Opus 4.8 pricing ($5/$25) while claiming near-Fable-5 quality, Gemini 3.6 Flash cuts output pricing 17%, and Grok's whole lineup turns out to carry a 2x long-context surcharge nobody's headline price mentions.

POCKET-35B

POCKET-35B

VIDRAFT quantizes its Darwin-36B-Opus MoE model into a 35B GGUF that runs on stock llama.cpp with no GPU, trading GPQA Diamond score for CPU and phone portability.

Claude Opus 5

Claude Opus 5

Anthropic's July 2026 release prices near-Fable-5 coding and agentic performance at Opus 4.8 rates, doubling Frontier-Bench scores and landing within 0.5 points of Fable 5 on CursorBench at half the cost.

SWE-1.7

SWE-1.7

Cognition's proprietary coding model powering Devin, scoring 42.3% on FrontierCode 1.1 Main at $1.97/task via Cerebras inference at 1000 tokens/sec.

Ling-3.0-flash

Ling-3.0-flash

InclusionAI's Ling-3.0-flash packs 124B parameters into a 5.1B-active hybrid-linear MoE that Ant Group claims matches its 1T flagship - but shipped with zero independently verifiable benchmark numbers.

Qwen3-VL-235B-A22B

Qwen3-VL-235B-A22B

Alibaba's flagship open-weight vision-language MoE beats every proprietary model on DocVQA at 96.5% and MathVista at 85.8%, but trails GPT-5.4 and Gemini 3.1 Pro on broad MMMU-Pro reasoning.

DeepSeek-VL2

DeepSeek-VL2

DeepSeek-VL2 is DeepSeek's open-weight Mixture-of-Experts vision-language model, activating just 4.5B of its 27B parameters to hit 93.3% on DocVQA and beat GPT-4o on OCRBench.

Qwen2.5-VL-72B-Instruct

Qwen2.5-VL-72B-Instruct

Alibaba's dense 72B vision-language model tops the open-weight DocVQA leaderboard at 96.4% and remains the default self-hosted choice for document and chart understanding.

Best AI for Document Understanding - July 2026

Best AI for Document Understanding - July 2026

Qwen3-VL-235B-A22B and Qwen2.5-VL-72B lead DocVQA above 96%, but the bigger story in July 2026 is that frontier labs have quietly stopped publishing comparable scores for their newest models.

AlayaWorld

AlayaWorld

AlayaWorld is a 15B open-weight video diffusion world model from Alaya Lab that sustains interactive, camera-controllable environments past 60 seconds.