
llm-d Joins CNCF - Kubernetes Gets a Native LLM Inference Stack
IBM Research, Red Hat, and Google Cloud donated llm-d to the CNCF at KubeCon EU, giving Kubernetes a production-grade distributed LLM inference framework built on vLLM.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

IBM Research, Red Hat, and Google Cloud donated llm-d to the CNCF at KubeCon EU, giving Kubernetes a production-grade distributed LLM inference framework built on vLLM.

The 15-person startup that launched an H100 GPU into space last November just became the fastest Y Combinator company ever to reach unicorn status.

OpenAI is shutting down its Sora video app and killing a $1B Disney deal as it pivots aggressively toward enterprise clients ahead of a public listing.

Apollo-owned Yahoo has launched Scout, an AI answer engine powered by Anthropic Claude, deploying it to 250 million US users as a direct challenge to Google, Perplexity, and ChatGPT.

GitHub Copilot inserts promotional tips for itself and Raycast into PR descriptions, with over 11,000 affected pull requests found across GitHub and GitLab.

HuggingFace's Transformers.js v4 rewrites its WebGPU runtime in C++, supports 200+ architectures, and delivers up to 4x faster inference in browsers and server-side JS runtimes.

Physical Intelligence is in talks to raise $1 billion at an $11 billion valuation, doubling in four months, as investors pour capital into AI software designed to run robots in the real world.

Mistral AI secures $830M in debt financing from seven banks to build a 13,800-GPU Nvidia GB300 cluster near Paris, targeting 200MW of European compute by 2027.

Google's Gemini 3.1 Flash Live beats GPT-4 Realtime 1.5 on Scale AI's Audio MultiChallenge and takes Search Live to 200+ countries - but it doesn't lead every benchmark.

NVIDIA's DSX Flex library and Emerald AI's Conductor platform let AI factories ramp GPU power up or down in seconds, unlocking faster grid connections and up to 100GW of new U.S. capacity.

Shopify activated Agentic Storefronts by default on March 24, making products from millions of merchants discoverable and purchasable inside ChatGPT, Microsoft Copilot, and Google's AI channels.

Meta releases SAM 3.1 with Object Multiplex, processing all tracked objects in one shared pass for 7x faster inference at 128 objects and improvements on 6 of 7 VOS benchmarks.