Recent Articles - Page 11

Latest News

Two World Models, One Multi-Agent Review Problem

Two World Models, One Multi-Agent Review Problem

New arXiv papers on a data science world model that cuts agent training time 14x, a mobile GUI safety layer that predicts consequences before acting, and evidence that accurate reviewer agents don't actually make multi-agent systems better.

View All News →

Guides

View All →

Reviews

View All →
Kimi K3 Review: Best at Code, Worse at Honesty

Kimi K3 Review: Best at Code, Worse at Honesty

Moonshot's Kimi K3 tops LMArena's Frontend Code Arena and undercuts Opus 4.8 on cost per task, but a tripled price tag, a rising hallucination rate, and an unresolved distillation question complicate the win.

Leaderboards

View All →

Models

View All →
Luma Ray3.2

Luma Ray3.2

Luma Ray3.2 is Luma AI's current flagship video model - native 16-bit HDR, 16-keyframe control, and the company's first full developer API, but still no native audio.

Pika 2.5

Pika 2.5

Pika Labs' flagship video model trades cinematic Elo rankings for the deepest creative-effects toolkit in AI video, plus a pivot into real-time agent video with PikaStream.

Haiper 2.x

Haiper 2.x

Haiper 2.x is the cheapest per-second AI video API on the market at $0.033/sec, now run by NetMind.AI after Haiper's consumer app shut down and its founders joined Microsoft.

Recent

AI Took 70% of Record $510B Venture Haul in H1

AI Took 70% of Record $510B Venture Haul in H1

Crunchbase data shows global startup investment hit $510 billion in H1 2026 - more than all of 2025 combined - with AI absorbing over 70% of Q2 capital and two labs capturing 43% of the total.

GPT-5.6 Sol Review: Strong Model, Thin Access

GPT-5.6 Sol Review: Strong Model, Thin Access

OpenAI's GPT-5.6 Sol tops Terminal-Bench 2.1 at 91.9% with its multi-agent Ultra mode, but reward-hacking findings and government-gated access keep it out of reach for nearly everyone.

LongCat-2.0

LongCat-2.0

Meituan's 1.6T-parameter open-source MoE coding model, trained end-to-end on 50,000 domestic Chinese ASICs, with native 1M token context and a 59.5 SWE-bench Pro score.

Science Agents, Jailbreak Defense, and Open-World Failures

Science Agents, Jailbreak Defense, and Open-World Failures

Three papers from today's arXiv: graph-native RL generates traceable scientific hypotheses, HARC defeats jailbreaks by coupling internal safety directions, and ICML 2026's OpenAgent shows how distributional shift breaks tool-use agents.

Gemini Omni Flash

Gemini Omni Flash

Google DeepMind's multimodal video generation model that creates 10-second clips with native audio from text, images, or video inputs - and lets you refine results through conversation.