PT-2026-93891 · Vllm · Vllm
CVE-2026-69147
·
Published
2026-09-16
·
Updated
2026-10-01
CVSS v3.1
6.5
Medium
| Vector | AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H |
Name of the Vulnerable Software and Affected Versions
vLLM versions prior to 0.28.0
Description
An issue exists where the engine fails to properly budget GPU memory when a user specifies a GPU decoder at request time. Specifically, request bodies for Chat Completions and Responses can set
media io kwargs.video.video backend to pynvvideocodec. The MediaConnector.fetch video() function forwards this choice to VideoMediaIO, allowing the use of the PyNvVideoCodec backend even if the startup configuration selected a software decoder.Because the
reserve mm ipc gpu memory() logic only allocates decoder memory based on static configuration, the request-selected backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that are not accounted for in the engine's KV-cache budget. An attacker submitting video requests to a deployment with PyNvVideoCodec installed can exhaust shared GPU memory, leading to request failures, worker crashes, or denial of service.Recommendations
Update vLLM to version 0.28.0 or later.
As a temporary mitigation, restrict the use of the
video backend parameter within media io kwargs to prevent users from selecting pynvvideocodec unless it was explicitly configured at startup.Exploit
Fix
Allocation of Resources Without Limits
Resource Exhaustion
Found an issue in the description? Have something to add? Feel free to write us 👾
Related Identifiers
Affected Products
Vllm