PT-2026-93891 · Vllm · Vllm

CVE-2026-69147

·

Published

2026-09-16

·

Updated

2026-10-01

CVSS v3.1

6.5

Medium

VectorAV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Name of the Vulnerable Software and Affected Versions vLLM versions prior to 0.28.0
Description An issue exists where the engine fails to properly budget GPU memory when a user specifies a GPU decoder at request time. Specifically, request bodies for Chat Completions and Responses can set media io kwargs.video.video backend to pynvvideocodec. The MediaConnector.fetch video() function forwards this choice to VideoMediaIO, allowing the use of the PyNvVideoCodec backend even if the startup configuration selected a software decoder.
Because the reserve mm ipc gpu memory() logic only allocates decoder memory based on static configuration, the request-selected backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that are not accounted for in the engine's KV-cache budget. An attacker submitting video requests to a deployment with PyNvVideoCodec installed can exhaust shared GPU memory, leading to request failures, worker crashes, or denial of service.
Recommendations Update vLLM to version 0.28.0 or later. As a temporary mitigation, restrict the use of the video backend parameter within media io kwargs to prevent users from selecting pynvvideocodec unless it was explicitly configured at startup.

Exploit

Fix

Allocation of Resources Without Limits

Resource Exhaustion

Found an issue in the description? Have something to add? Feel free to write us 👾

Weakness Enumeration

Related Identifiers

CVE-2026-69147
GHSA-8PW2-6JV3-MJ5J
PYSEC-2026-4178

Affected Products

Vllm