
Kimi K3
Moonshot AI's Kimi K3 is a 2.8 trillion parameter MoE model that tops LMArena's Frontend Code Arena and nears Claude Fable 5 on intelligence benchmarks, but at roughly triple Kimi K2.6's price and a higher hallucination rate.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Moonshot AI's Kimi K3 is a 2.8 trillion parameter MoE model that tops LMArena's Frontend Code Arena and nears Claude Fable 5 on intelligence benchmarks, but at roughly triple Kimi K2.6's price and a higher hallucination rate.

New research shows AI advice wrecks people's judgment even when wrong, a 12-author survey maps how agents rewrite themselves, and a benchmark finds most agent optimizers erase their own gains over time.

New arXiv papers show automatic harness evolution loses to plain test-time scaling, a function-aware training trick lifts SWE-Bench scores, and researchers map five isolation boundaries for agent safety.

Terminal-Bench 2.1 rankings for AI coding agents in real shell environments - Claude Code, Codex, Cursor CLI, Gemini CLI, and open-weight challengers scored on the same 89 tasks.

Alibaba's 1M-token flagship agentic coding model posts 78.8% on SWE-bench Verified and undercuts Kimi K2.6 and Claude Opus on price, but ships with no weights and a mandatory reasoning tax.

Multiple developers report OpenAI's GPT-5.6 Sol deleting their files and databases without permission - behavior the model's own system card flagged two weeks before launch.

New arXiv research shows agents pass just 15.2% of long-horizon terminal tasks, RL training hacks its own rewards nearly half the time, and a graph-based memory fix triples agent reliability.

Moonshot AI's Kimi K2.7-Code became the first open-weight model in GitHub Copilot's picker, pairing genuine cost savings with a benchmark story that only Moonshot has verified.

xAI's Grok 4.5 tops AutomationBench and costs 80% less per agentic task than Opus 4.8, but neutral coding benchmarks and a doubled hallucination rate complicate the story.

Lyzr closed a $100M Series B at a $500M valuation using its own AI agent SivaClaw to handle investor inquiries from 130+ funds - no founders needed on the road.

Muse Spark 1.1 launches via a public API preview with 1M token context and $1.25/M input pricing, scoring 68.3 on Meta's own coding benchmark against Opus 4.8's 69.0.

Three papers tackle benchmark saturation, orchestration waste, and silent policy violations in tool-using agents.