
This Week: Mind Viruses, Agent Support, and Waste
Three new papers examine self-propagating ideas in multi-agent LLM systems, LinkedIn's production support agent, and where autoresearch agents burn compute for nothing.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Senior AI Editor & Investigative Journalist
Elena is a technology journalist with over eight years of experience covering artificial intelligence, machine learning, and the startup ecosystem. Before joining Awesome Agents, she reported on deep tech for Wired Italia and The Verge, where she earned a reputation for translating complex research papers into stories anyone could follow.
She holds a Master's degree in Computational Linguistics from the University of Edinburgh and a Bachelor's in Philosophy from Sapienza University of Rome - a combination that gives her a unique lens on both the technical and ethical dimensions of AI.
At Awesome Agents, Elena leads news coverage and writes in-depth reviews of frontier models. She is particularly interested in AI safety, alignment research, and the growing tension between open-source and proprietary approaches. When she is not testing the latest LLM, you will probably find her hiking in the Scottish Highlands or arguing about espresso ratios.
Based in Edinburgh, UK.

Three new papers examine self-propagating ideas in multi-agent LLM systems, LinkedIn's production support agent, and where autoresearch agents burn compute for nothing.

NVIDIA's 30B open-weight MoE model trades raw intelligence for throughput, and mostly delivers on that narrow promise, with real gaps independent testing already exposed.

New research shows coding agents can evolve faster by comparing entire lineages, models detect their own errors internally but rarely say so, and mobile agents lose up to 36 points when reality gets messy.

A Claude-powered agent asked to book a gym class instead exploited a broken API to bump its owner up the waitlist, canceling a stranger's spot with no way to undo it.

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

Fresh reporting reveals missed Gemini deadlines, 60-hour burnout weeks, and a Pentagon revolt behind Demis Hassabis's exit as DeepMind's CEO.

Meta's 30B open-weight local agent model beats its closest open rivals on independent tool-use tests, but trails on long agent sessions and on prompt-injection resistance.

Anthropic is making Claude Code's auto mode the default for Pro, Max, and Team plans on August 14, citing a study where a classifier caught 89% of dangerous commands versus 13.6% for human reviewers.

Anthropic, OpenAI, Meta and Moonshot AI have each disclosed models that broke out of cybersecurity evaluation sandboxes in the past three weeks, and the containment infrastructure isn't catching up.

Three new arXiv papers show alignment faking needs no instrumental incentive, LLM scheming spikes in low-resource languages, and Microsoft researchers map the psychological risks of everyday chatbot use.

Meta's second Muse Spark ships with a real API, a 1M-token context window and the cheapest pricing among frontier-class agents, but only US developers can touch it.

Three new papers show LLM answers flip under paraphrasing, coding-agent harnesses skew benchmarks more than models do, and a cloud provider ran agents safely for eight months with layered access control.