Published event
Hardware
SecurityIncident
1 source(s)
v0.24.0
Summary
v0.24.0 vllm-project / vllm Public Uh oh! There was an error while loading.
Why it matters
This SecurityIncident is relevant to the technology intelligence record because it involves AMD, DeepSeek, Cohere, NVIDIA. The source article should remain the factual reference for follow-up coverage.
Key facts
- vllm-project / vllm Public Uh oh!
- There was an error while loading.
- Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.24.0 khluu released this 29 Jun 19:41 · 4088 commits to main since this release v0.24.0 ee0da84 vLLM v0.24.0 Release Notes Highlights This release features 571 commits from 256 contributors (77 new)!
- MiniMax-M3 : Added support for the new MiniMax-M3 model ( #45381 ), with a fast follow-on of BF16/FP8 indexer via MSA ( #45892 ), MXFP4 support ( #45896 ), FP8 sparse GQA ( #45744 ), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 ( #45725 ), fp8_per_channel for bf16 weights on MI300X ( #45854 ), FP8 KV-cache fix ( #45720 ), and packed-modules mapping ( #45794 ).
- A MiniMax-M2 perf regression was also fixed ( #45935 ).
- DeepSeek-V4 keeps maturing : Following its debut, DeepSeek-V4 received another large optimization pass — a FlashInfer sparse index cache (2–4% TTFT) ( #45863 ), prefill chunk-planning optimization (4% E2E throughput) ( #45061 ), a cluster-cooperative topK kernel for low-latency ( #43008 ), contiguous per-block KV allocations ( #44577 ), TEP=16 for the block-FP8 shared expert ( #46001 ), and native DSA indexer decode for next_n > 2 on SM100 ( #45322 ).
Entities in this story
Products
Docker→Related events