
MCP Drops Sticky Sessions to Scale Like the Web
The next Model Context Protocol spec removes session IDs and the initialize handshake entirely, letting MCP servers run behind ordinary round-robin load balancers for the first time.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

AI Infrastructure & Open Source Reporter
Sophie is a journalist and former systems engineer who covers AI infrastructure, open-source models, and the developer tooling ecosystem. She spent three years as a site reliability engineer at a cloud provider in Seattle before transitioning to tech journalism, which gives her writing an unusual level of technical depth - she understands distributed systems, GPU clusters, and inference optimization from the inside.
She studied Computer Engineering at the University of British Columbia and later completed a science communication fellowship at MIT. Her engineering background means she can read a model card, spot a misleading benchmark, and explain why quantization matters - all in the same paragraph.
At Awesome Agents, Sophie covers AI infrastructure news: new model releases, open-source launches, developer tools, deployment trends, and the hardware that makes it all run. She has a soft spot for underdog open-source projects that punch above their weight and a sharp eye for when a "breakthrough" is really just better marketing.
Based in Seattle, WA.

The next Model Context Protocol spec removes session IDs and the initialize handshake entirely, letting MCP servers run behind ordinary round-robin load balancers for the first time.

Google is reportedly building a chip line separate from its TPUs that hardwires parts of Gemini directly into silicon, promising up to 10x efficiency as a capacity crunch forces Cloud to turn away customers.

SK Group Chairman Chey Tae-won says customers want 60-100% more AI memory in 2027 than in 2026, and warns governments are starting to treat chip access as a matter of economic security.

Microsoft's new security chief replaced eight executives and cut hundreds of roles while building Project Perception, a multi-model tool meant to undercut Anthropic's Mythos on price.

PrismML compressed a 27B-parameter Qwen model from 54GB to under 4GB using 1-bit and ternary weights, and Apple is evaluating the technology for on-device Siri.

LM Studio launched Bionic, a standalone agent app that routes coding and document work between local open models and a Zero Data Retention cloud tier.

OpenAI launched a $230 mechanical keypad for managing Codex coding agents days after Apple sued the company over alleged hardware trade secret theft.

Governor Hochul signed an executive order pausing permits for data centers over 50 megawatts for up to a year, making New York the first US state to enact a statewide moratorium.

Nous Research is finalizing a round led by Robot Ventures and USV that would value the open-source Hermes agent maker at $1.5 billion, built on a training network that skips traditional data centers entirely.

Clem Delangue says cost is pushing companies off frontier APIs and onto open models. A16z's own CIO survey shows enterprise dollars still moving the other way.

Mistral's open-source Leanstral 1.5 scanned 57 repos and found five previously unreported bugs, including a silent integer overflow in a Rust zigzag decoder.

GPT-5.6 Sol is now live on Cerebras wafer-scale hardware at 750 tokens per second - roughly 10x faster than any GPU-based frontier model deployment in production.