
NVIDIA's Nemotron 3.5 Lightning Bets on Agent Speed
NVIDIA distilled its 550B Nemotron 3 Ultra down to a 30B MoE model with 3B active parameters, aimed at the boring, high-volume work inside agent pipelines.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

NVIDIA distilled its 550B Nemotron 3 Ultra down to a 30B MoE model with 3B active parameters, aimed at the boring, high-volume work inside agent pipelines.

Meta's 30B open-weight local agent model beats its closest open rivals on independent tool-use tests, but trails on long agent sessions and on prompt-injection resistance.

Meta's 30B open-weight agentic model distilled from Muse Spark runs on a single consumer GPU, ships under Apache 2.0, and leads Gemma4-31B and Qwen3.6-27B on 5 of 8 agentic benchmarks.

Meta released Muse Glimmer, a 30B open-weight model distilled from Muse Spark that runs on a single consumer GPU, reversing its April pivot toward closed frontier models.

Qwen3-30B-A3B is Alibaba's efficient MoE model that activates 3.3B of 30.5B parameters per token, matching much larger dense models on reasoning and agent benchmarks under Apache 2.0.

PrismML compressed a 27B-parameter Qwen model from 54GB to under 4GB using 1-bit and ternary weights, and Apple is evaluating the technology for on-device Siri.

LM Studio launched Bionic, a standalone agent app that routes coding and document work between local open models and a Zero Data Retention cloud tier.

Ollama raises $65M Series B led by Theory Ventures as 8.9 million monthly developers and 85% of Fortune 500 companies adopt local AI model deployment.

Adobe's deal to buy Topaz Labs is less about upscaling filters and more about NeuroStream, the on-device inference engine that runs large AI video models on consumer RTX GPUs.

Mistral AI's largest Ministral 3 model - 14B parameters, 256K context, Apache 2.0 license, multimodal, built for local deployment and agentic workflows.

Google DeepMind's new QAT checkpoints shrink the Gemma 4 E2B model to under 1GB, making serious on-device AI viable for phones and budget laptops.

Migrate from GPT-4o (now retired) or GPT-5.1 to self-hosted Llama 4 with near-zero code changes, but plan carefully for hardware, EU licensing, and realistic context window limits.