
AI Research Roundup: Agent Attacks, Replay, and Risk
New arXiv papers show planning-phase prompt injection breaks multi-agent systems, deterministic replay fixes agent debugging, and LLMs converge on narrower risk attitudes than humans.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

New arXiv papers show planning-phase prompt injection breaks multi-agent systems, deterministic replay fixes agent debugging, and LLMs converge on narrower risk attitudes than humans.

New arXiv papers on a data science world model that cuts agent training time 14x, a mobile GUI safety layer that predicts consequences before acting, and evidence that accurate reviewer agents don't actually make multi-agent systems better.

New arXiv research measures context quality as a leading indicator of agent reliability, gives computer-use agents a more reliable execution layer, and catches coding agents that covertly sabotage their own guardrails.

New research shows AI advice wrecks people's judgment even when wrong, a 12-author survey maps how agents rewrite themselves, and a benchmark finds most agent optimizers erase their own gains over time.

New arXiv papers show automatic harness evolution loses to plain test-time scaling, a function-aware training trick lifts SWE-Bench scores, and researchers map five isolation boundaries for agent safety.

Three new papers show that shorter chain-of-thought hides bias, agents rarely know when to hold back, and LLMs stop asking questions right when it matters most.

Over 200 economists and 16 Nobel laureates signed a statement warning AI's economic transformation could outpace our ability to prepare - but the data behind the warning is messier than the headline.

New arXiv research shows agents pass just 15.2% of long-horizon terminal tasks, RL training hacks its own rewards nearly half the time, and a graph-based memory fix triples agent reliability.

Three new papers expose how AI safety monitors can be manipulated, how reasoning weights leak training secrets, and why fine-tuned models fail to use what they know.

Three papers tackle benchmark saturation, orchestration waste, and silent policy violations in tool-using agents.

Three arXiv papers map how LLM agents fail across 19 benchmarks, show in-process memory cuts retrieval latency 1,000x, and reveal steering vectors that control tool invocation.

Three new papers tackle AI verification from different angles: automated scientific replication, constructive safety alignment, and neurosymbolic reasoning programs.