Articles Tagged "Diffusion Models"

AlayaWorld

AlayaWorld

AlayaWorld is a 15B open-weight video diffusion world model from Alaya Lab that sustains interactive, camera-controllable environments past 60 seconds.

Haiper 2.x

Haiper 2.x

Haiper 2.x is the cheapest per-second AI video API on the market at $0.033/sec, now run by NetMind.AI after Haiper's consumer app shut down and its founders joined Microsoft.

Mochi 1

Mochi 1

Genmo's Apache 2.0 licensed 10B-parameter video generator is the largest open-weight text-to-video model released, with no managed API and roughly $0.33 per clip to self-host on an H100.

SkyReels V4

SkyReels V4

SkyReels V4 is Skywork AI's unified multi-modal video model that jointly generates 1080p/32FPS video and synchronized audio from a single dual-stream diffusion transformer.

Runway Gen-4.5

Runway Gen-4.5

Runway's Gen-4.5 is a video generation model built on an Autoregressive-to-Diffusion architecture that held the top Artificial Analysis Elo position at launch with 1,247 points before Seedance 2.0 and Kling 3.0 surpassed it in early 2026.

DiffusionGemma 26B

DiffusionGemma 26B

DiffusionGemma 26B is Google DeepMind's open-weight discrete diffusion language model that generates 256 tokens in parallel, reaching 1,100+ tokens/sec on H100 - roughly 4x faster than autoregressive models of the same size.

NVIDIA SANA-WM

NVIDIA SANA-WM

NVIDIA's SANA-WM is a 2.6B-parameter hybrid linear diffusion transformer that generates 60-second 720p video with 6-DoF camera control on a single H100, built for embodied AI and robotics simulation.

NVIDIA SANA-WM - Minute-Scale Video on One GPU

NVIDIA SANA-WM - Minute-Scale Video on One GPU

NVIDIA NVLabs open-sourced SANA-WM, a 2.6B-parameter world model that generates 60-second 720p camera-controlled video on a single GPU, outperforming 14B+ competitors that need 8 GPUs.

Stable Audio 3.0 Ships Open Weights, 6-Min Songs

Stable Audio 3.0 Ships Open Weights, 6-Min Songs

Stability AI releases Stable Audio 3.0 as a four-model family with a new SAME autoencoder, open weights for three of four variants, and tracks up to 6 minutes 20 seconds - while Suno and Udio face ongoing copyright lawsuits over their training data.

HiDream-O1-Image

HiDream-O1-Image

HiDream-O1-Image is an 8B open-source text-to-image model with a pixel-space diffusion architecture that outperforms 32B FLUX.2 [dev] across five major benchmarks.