
Bonsai 1.7B
PrismML's smallest Bonsai checkpoint compresses Qwen3-1.7B under 0.5GB, and now powers the vision-language model behind Qualcomm's Snapdragon AR1 smart glasses demo.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

PrismML's smallest Bonsai checkpoint compresses Qwen3-1.7B under 0.5GB, and now powers the vision-language model behind Qualcomm's Snapdragon AR1 smart glasses demo.

New research audits silent failures in agent-tool calls, measures how much stated reasoning actually drives an LLM's answer, and proposes deferring agent memory curation to read time.

Alibaba's 2.4 trillion parameter flagship ships with real pricing and a published benchmark table, but the open-weight release it promised for this week still hasn't shown up.

Qwen3-30B-A3B is Alibaba's efficient MoE model that activates 3.3B of 30.5B parameters per token, matching much larger dense models on reasoning and agent benchmarks under Apache 2.0.

Three new arXiv papers show alignment faking needs no instrumental incentive, LLM scheming spikes in low-resource languages, and Microsoft researchers map the psychological risks of everyday chatbot use.

VIDRAFT quantizes its Darwin-36B-Opus MoE model into a 35B GGUF that runs on stock llama.cpp with no GPU, trading GPQA Diamond score for CPU and phone portability.

VIDRAFT compressed its leaderboard-climbing Darwin-36B-Opus into POCKET-35B, a GPU-free model for phones and CPUs, but its headline GPQA score depends on how you count.

Alibaba's flagship open-weight vision-language MoE beats every proprietary model on DocVQA at 96.5% and MathVista at 85.8%, but trails GPT-5.4 and Gemini 3.1 Pro on broad MMMU-Pro reasoning.

Alibaba's dense 72B vision-language model tops the open-weight DocVQA leaderboard at 96.4% and remains the default self-hosted choice for document and chart understanding.

Alibaba's 2.4 trillion parameter preview claims it trails only Claude Fable 5. I tested it for free at chat.qwen.ai and found a capable but slow model with zero benchmarks to back the claim.

Alibaba previewed a 2.4-trillion-parameter multimodal model at WAIC and said it ranks second only to Claude Fable 5, without publishing a single benchmark to back the claim.

Alibaba's 2.4 trillion parameter multimodal MoE model claims to trail only Claude Fable 5, but ships with no model card, no benchmark table, and no confirmed pricing.