Published event
Hardware
OpenSourceRelease
1 source(s)
v0.9.2
Summary
v0.9.2 vllm-project / vllm Public Uh oh! There was an error while loading.
Why it matters
This OpenSourceRelease is relevant to the technology intelligence record because it involves NVIDIA, Intel, AMD, OpenAI. The source article should remain the factual reference for follow-up coverage.
Key facts
- vllm-project / vllm Public Uh oh!
- There was an error while loading.
- Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.9.2 github-actions released this 07 Jul 17:05 · 14505 commits to main since this release v0.9.2 a5dd03c Highlights This release contains 452 commits from 167 contributors (31 new!) NOTE: This is the last version where V0 engine code and features stay intact.
- We highly recommend migrating to V1 engine.
- Engine Core Priority Scheduling is now implemented in V1 engine ( #19057 ), embedding models in V1 ( #16188 ), Mamba2 in V1 ( #19327 ).
- Full CUDA‑Graph execution is now available for all FlashAttention v3 (FA3) and FlashMLA paths, including prefix‑caching.
Entities in this story
Related events