Articles Tagged "Multi-Agent"

Model Steering, Angry Buyers, and Blind Judges

Model Steering, Angry Buyers, and Blind Judges

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

Two World Models, One Multi-Agent Review Problem

Two World Models, One Multi-Agent Review Problem

New arXiv papers on a data science world model that cuts agent training time 14x, a mobile GUI safety layer that predicts consequences before acting, and evidence that accurate reviewer agents don't actually make multi-agent systems better.

Sakana Fugu

Sakana Fugu

Sakana AI's orchestrator model that dynamically coordinates Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro to beat each of them individually on SWE-Bench Pro, GPQA-Diamond, and eight other benchmarks.