
Kimi K3 Tops Frontend Arena Just as Its Price Triples
Moonshot AI's Kimi K3 jumped 17 spots to #1 on LMArena's Frontend Code Arena, but the win comes with a tripled price tag and a weaker showing on broader intelligence benchmarks.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Moonshot AI's Kimi K3 jumped 17 spots to #1 on LMArena's Frontend Code Arena, but the win comes with a tripled price tag and a weaker showing on broader intelligence benchmarks.

Moonshot AI's Kimi K3 is a 2.8 trillion parameter MoE model that tops LMArena's Frontend Code Arena and nears Claude Fable 5 on intelligence benchmarks, but at roughly triple Kimi K2.6's price and a higher hallucination rate.

China's internet regulator approved Apple Intelligence for the local market, but only after Apple agreed to run Alibaba's Qwen and Baidu's models instead of its own.

Snowflake's reasoning-first text-to-SQL model tops the BIRD benchmark at 71.83% execution accuracy, trained with GRPO and a reward that only checks if the SQL runs correctly.

Mira Murati's Thinking Machines Lab released its first open-weight model, Inkling, and published benchmarks showing it losing to closed rivals on most of them.

Thinking Machines Lab's first open-weight model - a 975B-parameter MoE with native text, image, and audio reasoning, released under Apache 2.0 and tuned for customization on the Tinker platform.

Internet pioneer Vint Cerf has joined Innovation Labs to push DNSid, a DNS-anchored identity standard for AI agents, through the IETF after retiring from Google.

Nebius will sell Reflection AI over $1 billion in Nvidia GB300 compute through 2029, the open-source AI lab's second billion-dollar infrastructure deal in three weeks.

Nous Research's 36B open-weight model matches Hermes 4 70B on most benchmarks, tops RefusalBench on alignment, and is the first production model trained entirely on the Solana-secured Psyche network.

Nous Research is finalizing a round led by Robot Ventures and USV that would value the open-source Hermes agent maker at $1.5 billion, built on a training network that skips traditional data centers entirely.

Clem Delangue says cost is pushing companies off frontier APIs and onto open models. A16z's own CIO survey shows enterprise dollars still moving the other way.

Moonshot AI's Kimi K2.7-Code became the first open-weight model in GitHub Copilot's picker, pairing genuine cost savings with a benchmark story that only Moonshot has verified.