Flash attention

vLLM 0.17 Ships FlashAttention 4 and Live MoE Scaling

vLLM v0.17.0 adds FlashAttention 4, elastic expert parallelism for live MoE rescaling, full Qwen3.5 support, and a performance-mode flag, all in 699 commits from 272 contributors.

Flash attention

vLLM 0.17 Ships FlashAttention 4 and Live MoE Scaling

Google Analytics