Articles Tagged "AI Coding"

GPT-6 Astra Review: Genius Benchmarks, Thin Trust

GPT-6 Astra Review: Genius Benchmarks, Thin Trust

OpenAI's most capable model yet tops cybersecurity and agentic benchmarks, but a rocky rollout, user complaints of degraded output, and its own system card's warnings about hidden reasoning complicate the launch.

Qwen3.8-Max

Qwen3.8-Max

Alibaba's 2.4 trillion parameter flagship ships with real pricing and a published benchmark table, but the open-weight release it promised for this week still hasn't shown up.

Claude Code Drops Approval Prompts by Default

Claude Code Drops Approval Prompts by Default

Anthropic is making Claude Code's auto mode the default for Pro, Max, and Team plans on August 14, citing a study where a classifier caught 89% of dangerous commands versus 13.6% for human reviewers.

Muse Spark 1.1

Muse Spark 1.1

Meta's second Muse model ships a public API at $1.25/$4.25 per million tokens, a 1M-token context window, and the top score on Meta's own tool-use benchmarks.

Claude Opus 5

Claude Opus 5

Anthropic's July 2026 release prices near-Fable-5 coding and agentic performance at Opus 4.8 rates, doubling Frontier-Bench scores and landing within 0.5 points of Fable 5 on CursorBench at half the cost.