Recent Articles - Page 109

Latest News

Runway Builds a Model Router for AI Video and Audio

Runway Builds a Model Router for AI Video and Audio

Runway's new Media Router auto-selects the best video, image, or audio model for each API request by cost, quality, or speed - the first preference-based router built for generative media instead of text.

This Week in AI Research: Knowledge, Speed, Agent Risk

This Week in AI Research: Knowledge, Speed, Agent Risk

Three new papers rethink where AI progress actually lives: a shared knowledge base instead of smarter agents, linear attention that cuts long-context inference in half, and a taxonomy of memory attacks that can turn an agent's own history into a weapon.

View All News →

Guides

View All →

Reviews

View All →

Leaderboards

View All →

Models

View All →
Claude Opus 5

Claude Opus 5

Anthropic's July 2026 release prices near-Fable-5 coding and agentic performance at Opus 4.8 rates, doubling Frontier-Bench scores and landing within 0.5 points of Fable 5 on CursorBench at half the cost.

SWE-1.7

SWE-1.7

Cognition's proprietary coding model powering Devin, scoring 42.3% on FrontierCode 1.1 Main at $1.97/task via Cerebras inference at 1000 tokens/sec.

Ling-3.0-flash

Ling-3.0-flash

InclusionAI's Ling-3.0-flash packs 124B parameters into a 5.1B-active hybrid-linear MoE that Ant Group claims matches its 1T flagship - but shipped with zero independently verifiable benchmark numbers.

Recent

Best AI Data Analysis Tools in 2026

Best AI Data Analysis Tools in 2026

Compare the best AI data analysis tools of 2026 including Julius AI, ChatGPT Code Interpreter, and Claude analysis with pricing and features.

Best AI Meeting Assistants in 2026

Best AI Meeting Assistants in 2026

Compare the best AI meeting assistants of 2026 including Otter, Fireflies, Granola, and tl;dv with pricing, features, and recommendations.

CoT Control, Hidden Beliefs, and Dynamic Agent Benchmarks

CoT Control, Hidden Beliefs, and Dynamic Agent Benchmarks

New research shows reasoning models can't suppress their chain-of-thought, that they commit to answers internally long before their CoT reveals it, and that static benchmarks are inadequate for measuring real-world agent adaptability.