
Muse Glimmer
Meta's 30B open-weight agentic model distilled from Muse Spark runs on a single consumer GPU, ships under Apache 2.0, and leads Gemma4-31B and Qwen3.6-27B on 5 of 8 agentic benchmarks.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Meta's 30B open-weight agentic model distilled from Muse Spark runs on a single consumer GPU, ships under Apache 2.0, and leads Gemma4-31B and Qwen3.6-27B on 5 of 8 agentic benchmarks.

VIDRAFT quantizes its Darwin-36B-Opus MoE model into a 35B GGUF that runs on stock llama.cpp with no GPU, trading GPQA Diamond score for CPU and phone portability.

VIDRAFT compressed its leaderboard-climbing Darwin-36B-Opus into POCKET-35B, a GPU-free model for phones and CPUs, but its headline GPQA score depends on how you count.

Lunar Outpost and Firefly Aerospace will fly Nvidia Jetson modules on a lunar rover and orbiter this year, the first GPUs to run on and around the Moon.

PrismML compressed a 27B-parameter Qwen model from 54GB to under 4GB using 1-bit and ternary weights, and Apple is evaluating the technology for on-device Siri.

Bonsai 27B compresses Alibaba's Qwen3.6-27B into 1-bit and ternary weights, shrinking a 54GB model to as little as 3.9GB so it runs on an iPhone.

Mistral AI's mid-tier open-weight edge model - 8B parameters, 256K context, Apache 2.0 license, built for agentic pipelines and cost-sensitive production workloads.

Google DeepMind's new QAT checkpoints shrink the Gemma 4 E2B model to under 1GB, making serious on-device AI viable for phones and budget laptops.

NVIDIA Cosmos 3 is an open physical AI omnimodel with Mixture-of-Transformers architecture that natively handles text, images, video, sound, and robot actions in a single 16B or 64B model.

Mistral AI's smallest open-weight model - 3B parameters, 256K context, Apache 2.0 license, built for edge and cost-sensitive deployments.

South Korean startup LetinAR raises $18.5M to scale its PinTILT optical modules, which already power AI glasses and AR helmets as shipments hit 8.7 million units globally in 2025.

Google Chrome silently installs a 4 GB Gemini Nano model file on user devices with no consent prompt and re-downloads it if you delete it.