Articles Tagged "Multimodal"

Muse Glimmer

Muse Glimmer

Meta's 30B open-weight agentic model distilled from Muse Spark runs on a single consumer GPU, ships under Apache 2.0, and leads Gemma4-31B and Qwen3.6-27B on 5 of 8 agentic benchmarks.

Muse Spark 1.1

Muse Spark 1.1

Meta's second Muse model ships a public API at $1.25/$4.25 per million tokens, a 1M-token context window, and the top score on Meta's own tool-use benchmarks.

Runway Builds a Model Router for AI Video and Audio

Runway Builds a Model Router for AI Video and Audio

Runway's new Media Router auto-selects the best video, image, or audio model for each API request by cost, quality, or speed - the first preference-based router built for generative media instead of text.

Qwen3-VL-235B-A22B

Qwen3-VL-235B-A22B

Alibaba's flagship open-weight vision-language MoE beats every proprietary model on DocVQA at 96.5% and MathVista at 85.8%, but trails GPT-5.4 and Gemini 3.1 Pro on broad MMMU-Pro reasoning.

DeepSeek-VL2

DeepSeek-VL2

DeepSeek-VL2 is DeepSeek's open-weight Mixture-of-Experts vision-language model, activating just 4.5B of its 27B parameters to hit 93.3% on DocVQA and beat GPT-4o on OCRBench.

Qwen2.5-VL-72B-Instruct

Qwen2.5-VL-72B-Instruct

Alibaba's dense 72B vision-language model tops the open-weight DocVQA leaderboard at 96.4% and remains the default self-hosted choice for document and chart understanding.

Best AI for Document Understanding - July 2026

Best AI for Document Understanding - July 2026

Qwen3-VL-235B-A22B and Qwen2.5-VL-72B lead DocVQA above 96%, but the bigger story in July 2026 is that frontier labs have quietly stopped publishing comparable scores for their newest models.

Luma Ray3.2

Luma Ray3.2

Luma Ray3.2 is Luma AI's current flagship video model - native 16-bit HDR, 16-keyframe control, and the company's first full developer API, but still no native audio.

Hailuo 02

Hailuo 02

MiniMax's Hailuo 02 pairs a physics-aware architecture with some of the cheapest per-second video pricing in the industry, though newer rivals have since passed it on raw quality.

Qwen3.8-Max-Preview

Qwen3.8-Max-Preview

Alibaba's 2.4 trillion parameter multimodal MoE model claims to trail only Claude Fable 5, but ships with no model card, no benchmark table, and no confirmed pricing.

Bonsai 27B

Bonsai 27B

Bonsai 27B compresses Alibaba's Qwen3.6-27B into 1-bit and ternary weights, shrinking a 54GB model to as little as 3.9GB so it runs on an iPhone.