Articles Tagged "Agentic AI"

Claude Opus 5

Claude Opus 5

Anthropic's July 2026 release prices near-Fable-5 coding and agentic performance at Opus 4.8 rates, doubling Frontier-Bench scores and landing within 0.5 points of Fable 5 on CursorBench at half the cost.

SWE-1.7

SWE-1.7

Cognition's proprietary coding model powering Devin, scoring 42.3% on FrontierCode 1.1 Main at $1.97/task via Cerebras inference at 1000 tokens/sec.

Ling-3.0-flash

Ling-3.0-flash

InclusionAI's Ling-3.0-flash packs 124B parameters into a 5.1B-active hybrid-linear MoE that Ant Group claims matches its 1T flagship - but shipped with zero independently verifiable benchmark numbers.

Gemini 3.6 Flash

Gemini 3.6 Flash

Google DeepMind's workhorse Flash model cuts output tokens 17% versus Gemini 3.5 Flash, drops output pricing to $7.50/M, and cuts DeepSWE task tokens by 65% while trailing GPT-5.6 Luna and Grok 4.5 on raw coding scores.

Pika 2.5

Pika 2.5

Pika Labs' flagship video model trades cinematic Elo rankings for the deepest creative-effects toolkit in AI video, plus a pivot into real-time agent video with PikaStream.

Three Papers That Explain Why AI Agents Keep Failing

Three Papers That Explain Why AI Agents Keep Failing

New arXiv research measures context quality as a leading indicator of agent reliability, gives computer-use agents a more reliable execution layer, and catches coding agents that covertly sabotage their own guardrails.

Kimi K3

Kimi K3

Moonshot AI's Kimi K3 is a 2.8 trillion parameter MoE model that tops LMArena's Frontend Code Arena and nears Claude Fable 5 on intelligence benchmarks, but at roughly triple Kimi K2.6's price and a higher hallucination rate.