Published event
Hardware
Funding
1 source(s)
v0.28.0
Summary
v0.28.0 vllm-project / vllm Public Uh oh! There was an error while loading.
Why it matters
This Funding is relevant to the technology intelligence record because it involves DeepSeek, AMD, Meta, OpenAI. The source article should remain the factual reference for follow-up coverage.
Key facts
- vllm-project / vllm Public Uh oh!
- There was an error while loading.
- Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.28.0 khluu released this 26 Aug 09:46 · 1968 commits to main since this release v0.28.0 2cf0a69 v0.28.0 Highlights This release features 584 commits from 270 contributors (76 new)!
- Kimi-K3 performance push : a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support ( #50484 ), fused FlashKDA decode and prefill kernels ( #50654 , #51311 , #52458 ), SiTU activation support for MegaMoE ( #50510 ), GEMM-RS for sequence parallelism ( #52079 ), combined all-gathers with 1.5~3x kernel-level speedup ( #51070 ), an adaptive speculative token budget delivering ~60% better DSpark TTFT ( #51725 ), and optional shared-expert sharding saving ~17 GiB of memory per GPU ( #50912 ).
- Kimi-K3 also now runs on ROCm with the V2 model runner ( #51653 ).
- DeepSeek V4 : sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding ( #51538 ), joined by AMD Quark NVFP4 support ( #47972 ), reasoning-effort prompts and mappings ( #50580 ), sparse top-k metadata kernel optimizations ( #52084 , #51967 ), narrowed eager CUDA graph regions ( #51430 , #52401 ), and ROCm enablement on gfx11 and gfx950 ( #47017 , #52212 ).
Entities in this story
Related events