PT-2026-95164 · Vllm · Vllm
CVSS v4.0
8.7
High
| Vector | AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N |
Name of the Vulnerable Software and Affected Versions
vLLM versions prior to 0.29.1
Description
In prefill/decode disaggregated deployments, the software fails to properly clean up decode-side metadata for rejected inference requests. A remote attacker can submit multiple requests using the
max tokens parameter set to 0, leading to unbounded memory growth on the decode worker. This results in memory exhaustion and a denial of service (DoS) that persists until the worker restarts. Non-disaggregated deployments are not affected.Recommendations
Update vLLM to a version later than 0.29.0.
Avoid using the
max tokens parameter set to 0 in requests to the affected API endpoint until the update is applied.Exploit
Fix
Memory Leak
Found an issue in the description? Have something to add? Feel free to write us 👾
Weakness Enumeration
Related Identifiers
Affected Products
Vllm