Published event
ArtificialIntelligence Other 1 source(s)

v0.20.2

Updated September 26, 2026 · 2:46 PM · source date May 10, 2026

Summary

v0.20.2 vllm-project / vllm Public Uh oh! There was an error while loading.

Why it matters

This Other is relevant to the technology intelligence record because it involves DeepSeek, DeepSeek v4, gpt-oss, GPT. The source article should remain the factual reference for follow-up coverage.

Key facts
  • vllm-project / vllm Public Uh oh!
  • There was an error while loading.
  • Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.20.2 khluu released this 10 May 07:37 · 5903 commits to main since this release v0.20.2 bc150f5 vLLM v0.20.2 Highlights This release features 6 commits from 6 contributors (0 new)!
  • This is a small patch release with bug fixes for DeepSeek V4, gpt-oss, and Qwen3-VL Bug Fixes DeepSeek V4 sparse attention : Re-enable the persistent topk path on Hopper and ensure the memset kernel runs at CUDA graph capture time regardless of max_seq_len , fixing the MTP=1 hang on DeepSeek V4 ( #41665 , revert of #41605 ).
  • DeepSeek V4 KV cache : Fixed a "failure to allocate KV blocks" error in the V1 engine KV cache manager ( #41282 ).
  • gpt-oss MXFP4 + torch.compile : Plumbed hidden_dim_unpadded through the moe_forward fake op so MXFP4 works under torch.compile on v0.20.x ( #42002 , backport of #41646 ).
Entities in this story
Related events