PT-2026-99322 · Vllm · Vllm
CVSS v4.0
7.1
High
| Vector | AV:N/AC:L/AT:N/PR:L/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N |
Name of the Vulnerable Software and Affected Versions
vLLM versions prior to 0.29.0
Description
The software fails to enforce decoder prompt-length validation on the disaggregated serving endpoint '/inference/v1/generate'. When a request includes a 'features' multimodal payload, the system builds a multimodal EngineInput using caller-supplied
token ids without verifying them against model config.max model len. For specific multimodal processors that set skip prompt length check to true, such as Nemotron Parse, Whisper, and FireRedLID, the InputProcessor. validate prompt len() function returns immediately. This allows an overlong prompt to be processed as an EngineCoreRequest, which then causes a worker failure and denial of service when the input-batch is copied into a fixed-width NumPy row.Recommendations
Update to version 0.29.0.
Exploit
Fix
DoS
Resource Exhaustion
Found an issue in the description? Have something to add? Feel free to write us 👾
Weakness Enumeration
Related Identifiers
Affected Products
Vllm