Published event
Hardware HardwareLaunch 1 source(s)

v0.25.0

Updated September 26, 2026 · 2:46 PM · source date July 11, 2026

Summary

v0.25.0 vllm-project / vllm Public Uh oh! There was an error while loading.

Why it matters

This HardwareLaunch is relevant to the technology intelligence record because it involves OpenAI, DeepSeek, Mistral AI, NVIDIA. The source article should remain the factual reference for follow-up coverage.

Key facts
  • vllm-project / vllm Public Uh oh!
  • There was an error while loading.
  • Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.25.0 khluu released this 11 Jul 20:06 · 3529 commits to main since this release v0.25.0 702f481 vLLM v0.25.0 Release Notes Highlights This release features 558 commits from 232 contributors (64 new)!
  • Model Runner V2 is now the default for all dense models ( #44443 ).
  • Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS ( #46535 ), realtime embeddings ( #46762 ), prefix caching for Mamba hybrid models ( #42406 ), multimodal-prefix bidirectional attention ( #46942 ), and dynamic speculative decoding compatible with full CUDA graphs ( #45953 ).
  • PagedAttention has been removed ( #47361 ).
Entities in this story
Related events