Shared environment setup for every Cosmos3 cookbook (Reasoner and Generator). Each cookbook README links back here for the backend(s) it supports — pick the backend you want to run and follow that one section.
| Backend | Use it for | Used by |
|---|---|---|
| Cosmos Framework | Native PyTorch inference, launched with torchrun |
Reasoner, Generator (Audiovisual, Action, Transfer) |
| Diffusers | Direct generation with Cosmos3OmniPipeline |
Generator (Audiovisual) |
| TensorRT-LLM Generator | OpenAI-compatible VisualGen server (image/video/audio/action/transfer generation) | Generator (Audiovisual, Action, Transfer) |
| TensorRT-LLM Reasoner | OpenAI-compatible image/video reasoning server | Reasoner |
| Transformers | Hugging Face Transformers inference | Reasoner |
| vLLM | OpenAI-compatible reasoning server (image/video understanding) | Reasoner |
| vLLM-Omni | OpenAI-compatible generation server (image/video/audio/action/transfer) | Generator (Audiovisual, Action, Transfer) |
| Cosmos3 Certified NIM | Prebuilt NGC container serving either the Generator (image/video/action/transfer) or the Reasoner from one image | Reasoner, Generator (Audiovisual, Action, Transfer) |
| Reasoner NIM | Prebuilt OpenAI-compatible reasoning server (image/video understanding); no venv | Reasoner |
| Generator NIM | Prebuilt NGC container serving the Cosmos3 Generator for Text-to-Video and Image-to-Video inference | Generator (Audiovisual) |
-
Linux with NVIDIA GPU access.
-
uv,git, andgit-lfsinstalled. -
Hugging Face access to the gated Cosmos3 model repos. Generator also requires access to the gated nvidia/Cosmos-1.0-Guardrail HF repository. Authenticate once before the first run:
uvx hf@latest auth login # or: export HF_TOKEN=<your_token>
To disable the guardrail, set
enable_safety_checker=False(Diffusers),TRTLLM_DISABLE_COSMOS3_GUARDRAILS=1oruse_guardrails: falsethroughextra_params(TensorRT-LLM),guardrails: false(vLLM-Omniextra_params/extra_args), or--no-guardrails(Cosmos Framework). For Generator NIM set environment variablesNIM_ENABLE_TEXT_GUARDRAILS=0 NIM_ENABLE_VIDEO_GUARDRAILS=0. -
NIMs don't need Hugging Face access; instead, an NGC API key is required (used as
NGC_API_KEY). You can generate one on build.nvidia.com or NGC, then rundocker login nvcr.ioonce (username$oauthtoken, password = your key). This repository uses the Reasoner NIM imagenvcr.io/nim/nvidia/cosmos3-reasonerand the Generator NIM imagenvcr.io/nim/nvidia/cosmos3-generator. -
For the Cosmos Framework backend: access to
git@github.com:NVIDIA/cosmos-framework.git. -
Enough local disk for the venv/image, the uv cache, and the model cache. Nano downloads plus CUDA dependencies can take tens of GiB.
Several backends pin a CUDA build of torch/vllm that must match your NVIDIA
driver. Pick the tag that matches the CUDA version your driver supports:
| Driver CUDA | Backend tag | Notes |
|---|---|---|
| 13.x | cu130 |
Default in the notebooks. |
| 12.x | cu128 |
Use when a compatible wheel is available for the selected package version. |
vLLM does not publish a wheel for every CUDA minor version, so
--torch-backend=auto is not reliable here — choose the pair that matches your
driver.
Native PyTorch inference through the Cosmos Framework checkout. Used by the
run_*_with_cosmos_framework.ipynb notebooks and the Cosmos Framework
quickstarts.
From the cosmos repo root, clone (or reuse) the framework checkout:
mkdir -p packages
git clone https://github.com/NVIDIA/cosmos-framework.git packages/cosmos3
cd packages/cosmos3Install the framework dependencies into its venv. The inference path currently
imports modules from the training extras, so use the *-train dependency group
that matches your driver (see CUDA driver and the cuXXX backend):
# lerobot tracks test artifacts with git-LFS that this cookbook does not need;
# skipping smudge avoids failures from missing LFS blobs in uv's git mirror.
export GIT_LFS_SKIP_SMUDGE=1
# CUDA 13 driver (default):
uv sync --all-extras --group=cu130-train
# CUDA 12.x driver:
# uv sync --all-extras --group=cu128-trainThe notebooks honor COSMOS3_UV_GROUP (default cu130-train); set
export COSMOS3_UV_GROUP=cu128-train before launching them on CUDA 12.x systems.
This produces a venv at packages/cosmos3/.venv. Run framework commands either
by activating it (source .venv/bin/activate) or via its absolute interpreter
(.venv/bin/python, .venv/bin/torchrun).
For CUDA 13, NVIDIA documents the NGC PyTorch container
nvcr.io/nvidia/pytorch:25.09-py3 as the recommended starting point; for CUDA 12 use
nvcr.io/nvidia/pytorch:25.06-py3. See
Cosmos Framework setup.
Inside that image (or any minimal GPU host), install the system packages below before
your first torchrun inference — uv sync --all-extras alone is not enough for
guardrails.
Framework inference enables guardrails by default. The video guardrail path imports OpenCV (via RetinaFace), which needs graphics libraries that are often missing on headless servers and minimal containers.
From packages/cosmos3 (or the framework repo root), with apt-get available.
NGC and many training containers run as root — use apt-get directly (no sudo).
On a normal host where you are not root, prefix with sudo.
apt-get update
apt-get install -y --no-install-recommends \
curl ffmpeg git-lfs libgl1 libglib2.0-0 libx11-dev libxcb1 tree wgetVerify OpenCV imports after source .venv/bin/activate:
python -c "import cv2; print(cv2.__version__)"If you see libxcb.so.1: cannot open shared object file, the libxcb1 / libgl1
packages above were not installed.
When using the NGC PyTorch base image, clear LD_LIBRARY_PATH after activating the
venv so the container’s bundled libtorch does not shadow the venv (see
Cosmos Framework FAQ — PyTorch import inside NGC):
source .venv/bin/activate
export LD_LIBRARY_PATH=Guardrails also require Hugging Face access to the gated safety models (accept the
license and set HF_TOKEN as in Prerequisites). To disable guardrails
for a one-off run, pass --no-guardrails to cosmos_framework.scripts.inference.
Direct generation with Cosmos3OmniPipeline (Generator · Audiovisual). Create a
venv and install the backend, choosing --torch-backend to match your driver
(see CUDA driver and the cuXXX backend):
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install --torch-backend=cu130 \
"diffusers @ git+https://github.com/huggingface/diffusers.git" \
accelerate \
av \
cosmos_guardrail \
huggingface_hub \
imageio \
imageio-ffmpeg \
torch \
torchvision \
transformersOpenAI-compatible VisualGen server for Generator audiovisual text-to-image,
text-to-video, image-to-video, video-to-video, synchronized audio, Transfer, and
Action examples.
Initial Cosmos3 support was added in TensorRT-LLM PR
#14824, synchronized audio
in #14827, and
video-to-video in #16155,
Transfer in #16394, and
Action in #17325.
The DMD2-distilled four-step checkpoints were added in
#16563 (text-to-image) and
#16690 (image-to-video), and
Cosmos3-Edge (Nemotron-dense backbone) in
#16773.
These changes are merged on TensorRT-LLM main. The Action and Transfer
notebooks were executed against source revision
bca6761ab84fbcd58fc7f914eade7de48b32e35e.
Use that revision to reproduce their request contract, or a newer build with
the same API.
The source revision is significant: a package version of 1.3.0rc26 alone
does not establish compatibility with the action image-decoding path.
Install TensorRT-LLM following its upstream documentation.
To build TensorRT-LLM from source, follow NVIDIA's Build from Source guide. This is the right path when you need a checkout that contains a recent Cosmos3 VisualGen change before it is available in your installed package or release image.
apt-get update && apt-get -y install git git-lfs
git lfs install
git clone https://github.com/NVIDIA/TensorRT-LLM.git
cd TensorRT-LLM
# Source revision used for the Action/Transfer notebook validation below.
git checkout bca6761ab84fbcd58fc7f914eade7de48b32e35e
git submodule update --init --recursive
git lfs pull
# Pick a devel tag from the upstream build-from-source guide or NGC.
docker pull nvcr.io/nvidia/tensorrt-llm/devel:<tag>
docker run --rm -it \
--ipc=host \
--ulimit memlock=-1 --ulimit stack=67108864 \
--gpus=all \
--volume "$PWD":"$PWD" \
--workdir "$PWD" \
nvcr.io/nvidia/tensorrt-llm/devel:<tag>
# Inside the container:
apt-get update && apt-get -y install ffmpeg
python3 scripts/build_wheel.py --use_ccache --skip_building_wheel --linking_install_binary
pip install -e .For Python-only changes, the upstream guide also documents
TRTLLM_USE_PRECOMPILED=1 pip install -e . to reuse precompiled binaries while
installing the checkout in editable mode.
Then install the Cosmos3 guardrail package in the same environment unless you explicitly disable guardrails before starting the server:
pip install cosmos_guardrail==0.3.0
# On headless servers without libGL.so.1, replace the OpenCV wheel pulled in by
# cosmos_guardrail with the matching headless build:
pip uninstall -y opencv-python
pip install opencv-python-headless==5.0.0.93With cosmos_guardrail==0.3.0, NLTK's path checks can reject tokenizer or
dictionary files that are symlinks from a Hugging Face snapshot into its
blobs/ directory. Prepare a separate copy of the NLTK data as regular files
on the server, in the same container and shell used to launch TensorRT-LLM:
export NLTK_DATA="$(mktemp -d "${TMPDIR:-/tmp}/cosmos3-nltk.XXXXXX")"
python3 - <<'PY'
import os
import shutil
from pathlib import Path
import nltk
from huggingface_hub import snapshot_download
snapshot = snapshot_download(
"nvidia/Cosmos-1.0-Guardrail",
revision="cf03c0395fac8c4de386c0bdab12cc4fc8d66362",
allow_patterns=["blocklist/**"],
)
source = Path(snapshot) / "blocklist" / "nltk_data"
destination = Path(os.environ["NLTK_DATA"]).resolve()
shutil.copytree(source, destination, symlinks=False, dirs_exist_ok=True)
assert not any(path.is_symlink() for path in destination.rglob("*"))
# Check both resource lookups used by the text blocklist before starting a GPU server.
nltk.data.path[:] = [str(destination)]
tokens = nltk.word_tokenize("You are an autonomous vehicle planning system.")
assert nltk.WordNetLemmatizer().lemmatize("vehicles") == "vehicle"
print("Guardrail NLTK data ready:", destination, tokens)
PYKeep NLTK_DATA exported when starting the server below. Repeat this setup
after recreating the container or removing the temporary directory. Running it
only in the client notebook's environment does not configure a remote server.
This workaround does not rewrite cached files or symlinks, and keeps NLTK
path security and use_guardrails=True enabled.
The separate No safety models found, returning safe warning in guardrail
0.3.0 refers to its intentionally empty video-content classifier list. Text
checks and face blurring remain configured; the NLTK workaround does not enable
video-content classification or suppress that warning.
Set the TensorRT-LLM source root for the shared VisualGen config YAMLs. Run this
from inside the TensorRT-LLM checkout — the directory the git clone above
created, which is where examples/ lives — or point TRTLLM_ROOT at that
checkout explicitly. trtllm-serve only reports a bad --visual_gen_args path
after it has started, so check it here instead:
export TRTLLM_ROOT="${TRTLLM_ROOT:-$PWD}"
export COSMOS3_TRTLLM_PORT="${COSMOS3_TRTLLM_PORT:-8000}"
test -d "$TRTLLM_ROOT/examples/visual_gen/configs" \
|| echo "TRTLLM_ROOT=$TRTLLM_ROOT does not look like a TensorRT-LLM checkout"Cosmos3-Nano (single GPU):
trtllm-serve nvidia/Cosmos3-Nano \
--visual_gen_args "$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml" \
--port "$COSMOS3_TRTLLM_PORT"Cosmos3-Super (four GPUs; CFG parallelism with Ulysses, plus parallel VAE):
torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \
nvidia/Cosmos3-Super \
--visual_gen_args "$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-super-4gpu.yaml" \
--port "$COSMOS3_TRTLLM_PORT"Cosmos3-Edge (single GPU):
trtllm-serve nvidia/Cosmos3-Edge \
--port "$COSMOS3_TRTLLM_PORT"Edge is the compact 4B checkpoint. Its 480p-native generation defaults
(832x480 with 121 frames, 50 UniPC steps on the checkpoint-declared native flow
schedule, guidance 5.0, flow shift 3.0) are read from the checkpoint, so it takes
no --visual_gen_args override. Edge text-to-image goes to
/v1/images/generations with "output_type": "image" in extra_params (video
mode would otherwise apply Cosmos3's video negative prompt to a still); the two
video modes go to /v1/videos/generations. TensorRT-LLM serves Edge for
text-to-image, text-to-video, and image-to-video only: Edge has no audio tower, its action
weights are not served by this pipeline, and video-to-video is validated for Nano
and Super. Requests outside the model card's validated envelope (256p/480p,
50-150 frames, 12-30 FPS) still run and log an advisory line.
Cosmos3-Super-Text2Image-4Step (single GPU; DMD2-distilled text-to-image):
trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \
--visual_gen_args "$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-t2i-1gpu.yaml" \
--port "$COSMOS3_TRTLLM_PORT"Cosmos3-Super-Image2Video-4Step (single GPU; DMD2-distilled image-to-video):
trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step \
--port "$COSMOS3_TRTLLM_PORT"Both distilled students run a fixed four-step stochastic schedule read from the
checkpoint's scheduler config, with classifier-free guidance baked into the
weights. TensorRT-LLM supplies both values and rejects a request that sends a
different num_inference_steps, or a guidance_scale other than 1.0, so leave
both out of the request. Cosmos3-Super-Image2Video-4Step also declares
default_use_system_prompt: true, which applies only while the request leaves
use_system_prompt unset. The text-to-image student deploys at 1024x1024, the
shape cosmos3-t2i-1gpu.yaml warms; the image-to-video student deploys at the
default 720p x 189-frame omni shape and needs no config file. The
distilled 4-step notebook
runs both against a running server. These students cover text-to-image and
image-to-video only; use the base checkpoints for text-to-video, video-to-video,
and synchronized audio.
The server exposes /health, the blocking /v1/videos/sync, the asynchronous
/v1/videos, and /v1/images/generations. The older
/v1/videos/generations spelling is a deprecated alias of /v1/videos/sync.
The base audiovisual notebook uses /v1/images/generations for text-to-image
and /v1/videos/sync for text-to-video, image-to-video, video-to-video, and
synchronized audio. Text-to-image sets extra_params.output_type="image" and
returns a base64-encoded PNG. Image-to-video uploads multipart
image_reference; video-to-video uses video_reference. Synchronized audio is
enabled with enable_audio: true in extra_params and is muxed into the output
video.
Every video request explicitly selects MP4. Keep ffmpeg on the server PATH:
without it, the request fails early instead of returning browser-incompatible
AVI or dropping generated audio. Requests send
Cosmos3 controls through extra_params, so use a TensorRT-LLM build that includes
the Cosmos3 VisualGen API schema. The notebook sets request-level
max_sequence_length=4096 for longer structured JSON prompts.
Transfer uses the synchronous /v1/videos/sync route. For server-derived edge
or blur, upload the raw source video as multipart video_reference and set the
corresponding extra_params hint to true. For a precomputed edge, blur,
depth, segmentation, or WSM control, base64-encode the control inside its hint;
no top-level reference is needed. The server decodes inline media to bytes at the
HTTP boundary. TensorRT-LLM uses use_guardrails for its per-request safety
switch; guardrails, control_path, and other vLLM-Omni-only names are not
interchangeable.
Action requests use the same synchronous route and upload an image as
image_reference or a video as video_reference. For the checked-in AV
examples, use the Cosmos Framework reference prompt shown in the Action
cookbook; current TensorRT-LLM ignores the legacy view_point field.
Because an action trajectory cannot be represented in MP4 or
AVI, format=auto resolves to safetensors; the payload contains named video,
action, and frame_rate tensors. The asynchronous /v1/videos route also
supports this payload: poll GET /v1/videos/{id}, then download it from
GET /v1/videos/{id}/content.
OpenAI-compatible reasoning server for image and video understanding. Run
TensorRT-LLM in the prebuilt
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc22
container, or follow the
TensorRT-LLM Generator source-build instructions to
install TensorRT-LLM from source.
Authenticate with Hugging Face before loading gated checkpoints.
Install headless OpenCV in the environment that runs the server. Cosmos3 Reasoner uses it to process vision inputs:
python -m pip install opencv-python-headlessRun one of these commands inside the release container or activated local environment.
Cosmos3-Nano (single GPU, port 8001):
CUDA_VISIBLE_DEVICES=0 \
trtllm-serve nvidia/Cosmos3-Nano \
--host 0.0.0.0 \
--port 8001 \
--max_num_tokens 32768Cosmos3-Super (four GPUs, port 8001):
CUDA_VISIBLE_DEVICES=0,1,2,3 \
trtllm-serve nvidia/Cosmos3-Super \
--host 0.0.0.0 \
--port 8001 \
--tensor_parallel_size 4 \
--max_num_tokens 32768The server exposes /health and the OpenAI-compatible API at
http://localhost:8001/v1. See the
Reasoner TensorRT-LLM notebook for image
and video requests.
Local Python inference for the Cosmos3 Reasoner. This backend uses the Transformers Cosmos3 integration and loads only the Reasoner tower.
Nano / Super load the Reasoner tower from the unified
nvidia/Cosmos3-Nano or nvidia/Cosmos3-Super checkpoint with
Cosmos3OmniForConditionalGeneration. Support first appears in Transformers
v5.11.0. Create a venv and install Transformers 5.11.0 or newer:
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install --torch-backend=auto \
accelerate \
av \
pillow \
"safetensors>=0.8.0" \
torch \
"torchvision==0.25.0" \
"transformers>=5.11.0"--torch-backend=auto lets uv detect the CUDA build of torch/torchvision
that matches your NVIDIA driver. Pin a backend such as cu128 or cu130 if
your environment needs an explicit CUDA wheel.
Cosmos3-Edge uses a separate Transformers integration
(AutoModelForImageTextToText / Cosmos3EdgeForConditionalGeneration) with
nvidia/Cosmos3-Edge. Do not load Edge with Cosmos3OmniForConditionalGeneration.
Edge support is on Transformers main and is not yet in a stable PyPI release;
install from GitHub instead of the
transformers>=5.11.0 pin above:
uv pip install "transformers @ git+https://github.com/huggingface/transformers.git"See the Reasoner Transformers quickstart for runnable Nano, Super, and Edge image examples and video input notes.
OpenAI-compatible reasoning server for the Reasoner cookbook (image/video
understanding). Native Cosmos3 Reasoner support first appears in the vLLM
v0.23.0 stable release:
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
# CUDA 13 driver:
uv pip install --torch-backend=cu130 "vllm>=0.23.0"
# CUDA 12.x driver:
# uv pip install --torch-backend=cu128 "vllm>=0.23.0"The vLLM wheel and the torch backend must be compatible — see
CUDA driver and the cuXXX backend.
If your vLLM build reports that DeepGEMM is unavailable, disable it before starting the server:
export VLLM_USE_DEEP_GEMM=0When launching with
.venv/bin/vllminstead of activating the venv, make sure.venv/binis onPATH(e.g.source .venv/bin/activate). FlashInfer's just-in-time kernel build shells out toninja, which lives in the venv.
All Reasoner cookbooks talk to an OpenAI-compatible chat-completions API. After
installing vLLM, run the commands below from
cookbooks/cosmos3/reasoner (same working directory as
run_with_vllm.ipynb). That sets
$(dirname "$(pwd)") to <cosmos>/cookbooks/cosmos3, which matches the
notebook's COSMOS3_MEDIA_ROOT.
Cosmos3-Nano (single GPU, port 8000):
CUDA_VISIBLE_DEVICES=0 \
vllm serve nvidia/Cosmos3-Nano \
--tensor-parallel-size 1 \
--mm-encoder-tp-mode data \
--async-scheduling \
--allowed-local-media-path "$(dirname "$(pwd)")" \
--media-io-kwargs '{"video": {"num_frames": -1}}' \
--port 8000Cosmos3-Super (four GPUs; default in run_with_vllm.ipynb, port 8001):
export COSMOS3_MEDIA_ROOT="$(dirname "$(pwd)")"
export VLLM_PORT="${VLLM_PORT:-8001}"
CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--mm-encoder-tp-mode data \
--async-scheduling \
--allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
--media-io-kwargs '{"video": {"num_frames": -1}}' \
--port "$VLLM_PORT"Cosmos3-Edge (single GPU, port 8000):
CUDA_VISIBLE_DEVICES=0 \
vllm serve nvidia/Cosmos3-Edge \
--tensor-parallel-size 1 \
--mm-encoder-tp-mode data \
--async-scheduling \
--allowed-local-media-path "$(dirname "$(pwd)")" \
--media-io-kwargs '{"video": {"num_frames": -1}}' \
--port 8000The Super notebook polls /health for up to 1800 seconds on first start while CUDA
graphs compile.
| Option | Use |
|---|---|
--tensor-parallel-size |
Number of GPUs for tensor-parallel inference |
--mm-encoder-tp-mode data |
Data parallelism for the visual encoder |
--media-io-kwargs '{"video": {"num_frames": -1}}' |
Lets the processor see all frames before downstream sampling |
--allowed-local-media-path |
Must cover local file:// media paths; defaults to <cosmos>/cookbooks/cosmos3 when run from cookbooks/cosmos3/reasoner |
OpenAI-compatible generation server (image/video/audio/action/transfer) for the Generator cookbooks.
Cosmos3 checkpoints can exceed the default server init timeout — always pass
--init-timeout 1800 on every vllm serve command below.
The vLLM-Omni server loads gated
nvidia/Cosmos-1.0-Guardrail
at startup by default. Without Hugging Face access to that repo, the server
exits before serving requests. Per-request guardrails: false in extra_params
(see Prerequisites) does not fix this — the guardrail models
must load at startup.
To disable guardrails server-wide (you are responsible for
license compliance),
add --no-guardrails to any vllm serve command below:
vllm serve nvidia/Cosmos3-Nano \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--no-guardrails \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800See also the vLLM-Omni Cosmos3-Nano recipe.
The released image vllm/vllm-omni:cosmos3 supports the Generator modalities,
including transfer controls. Pull once:
docker pull vllm/vllm-omni:cosmos3Set paths once; adjust for your checkout and cache location:
export HF_HOME="${HF_HOME:-$HOME/.cache/huggingface}"
export COSMOS3_WORKDIR="${COSMOS3_WORKDIR:-$(pwd)}"
export COSMOS3_HOST_PORT="${COSMOS3_HOST_PORT:-8000}"The container listens on port 8000; -p "${COSMOS3_HOST_PORT}:8000" publishes it
on the host. Generator notebooks often use COSMOS3_HOST_PORT=8001 so port 8000
stays free for a Reasoner server. The Docker commands run from /workspace, so
repo-local paths such as cookbooks/... resolve inside the container.
Cosmos3-Nano (single GPU):
docker run --runtime nvidia --gpus '"device=0"' \
-e CUDA_DEVICE_ORDER=PCI_BUS_ID \
-v "${HF_HOME}:/root/.cache/huggingface" \
-v "${COSMOS3_WORKDIR}:/workspace" \
-p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
-w /workspace \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Nano \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800Cosmos3-Super (all GPUs; add tensor parallelism and layerwise offload):
docker run --runtime nvidia --gpus all \
-v "${HF_HOME}:/root/.cache/huggingface" \
-v "${COSMOS3_WORKDIR}:/workspace" \
-p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
-w /workspace \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Super \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--port 8000 \
--init-timeout 1800Cosmos3-Edge (single GPU):
docker run --runtime nvidia --gpus all \
-v "${HF_HOME}:/root/.cache/huggingface" \
-v "${COSMOS3_WORKDIR}:/workspace" \
-p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Edge \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800Cosmos3-Super-Text2Image-4Step (all GPUs; add tensor parallelism and layerwise offload):
docker run --runtime nvidia --gpus all \
-v "${HF_HOME}:/root/.cache/huggingface" \
-v "${COSMOS3_WORKDIR}:/workspace" \
-p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Super-Text2Image-4Step \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--port 8000 \
--init-timeout 1800Cosmos3-Super-Image2Video-4Step (all GPUs; add tensor parallelism and layerwise offload):
docker run --runtime nvidia --gpus all \
-v "${HF_HOME}:/root/.cache/huggingface" \
-v "${COSMOS3_WORKDIR}:/workspace" \
-p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Super-Image2Video-4Step \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--port 8000 \
--init-timeout 1800Mount any directory that holds local media or action JSON files referenced in
requests. Set --allowed-local-media-path / (as above) when the whole container
filesystem should be readable.
vLLM-Omni prints Application startup complete. when the API is ready.
To install from main instead of Docker, create a venv and pick the CUDA build
that matches your driver (see
CUDA driver and the cuXXX backend):
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
# CUDA 13 driver:
uv pip install --torch-backend=cu130 \
"vllm-omni @ git+https://github.com/vllm-project/vllm-omni.git@main"
# CUDA 12.x driver:
# uv pip install --torch-backend=cu128 \
# "vllm-omni @ git+https://github.com/vllm-project/vllm-omni.git@main"Run the same vllm serve arguments as in the Docker commands above, directly on
the host (no docker run wrapper):
vllm serve nvidia/Cosmos3-Nano \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800For Super, add --tensor-parallel-size 4 --enable-layerwise-offload.
Additional parallelism options (Docker or native):
| Option | Use |
|---|---|
--cfg-parallel-size 2 |
Runs positive and negative CFG branches on two GPUs |
--ulysses-degree 2 |
Ulysses sequence parallelism across GPUs |
Ensure the server has enough GPUs for the product of enabled degrees
(tensor_parallel_size × cfg_parallel_size × ulysses_degree).
Prebuilt NGC containers for Cosmos3. Like vLLM-Omni, NIM runs from Docker, so
there is no venv or --torch-backend to manage. Unlike the Hugging Face based
backends, NIM authenticates with an NGC API key instead of a Hugging Face token
(see Prerequisites).
Authenticate Docker to NGC once:
export NGC_API_KEY=<your_key>
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdinEach NIM below exposes readiness at GET /v1/health/ready after model
download, engine initialization, and warmup complete.
A single prebuilt container that serves either runtime — one runtime per
container, chosen at launch with NIM_MODEL_TYPE:
- Generator (
POST /v1/infer) — text-to-image, text-to-video, image-to-video, video-to-video, Action (forward dynamics, inverse dynamics, policy), and Transfer. Each request sets an explicit top-levelmodel_mode. - Reasoner — OpenAI-compatible image and video understanding through Chat Completions and the Responses API.
Start the Nano Generator runtime:
export NGC_API_KEY=<your_key>
export LOCAL_NIM_CACHE="${LOCAL_NIM_CACHE:-$HOME/.cache/nim}"
mkdir -p "$LOCAL_NIM_CACHE"
docker run -d --name cosmos3-generator \
--gpus '"device=0"' \
--shm-size 16g \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
--ulimit nofile=65536:65536 \
-p 8000:8000 \
-e NGC_API_KEY \
-e NIM_MODEL_TYPE=generator \
-e NIM_MODEL_VARIANT=nano \
-e NIM_PERF_PROFILE=latency \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
nvcr.io/nim/nvidia/cosmos3:2.0.0Set NIM_MODEL_TYPE=reasoner for the Reasoner runtime. NIM_MODEL_VARIANT
selects the checkpoint: nano or super for either runtime, plus
nano-droid, super-t2i, super-t2i-4step, super-i2v, and
super-i2v-4step for the Generator. On DGX Spark/GB10 and Jetson AGX Thor, the
Reasoner also needs NIM_GPU_MEMORY_UTILIZATION set explicitly.
The Cosmos3 Certified NIM cookbook is the full guide — deployment, support matrix, configuration, and API reference, with worked examples for generation, action, transfer, and reasoning.
A prebuilt container that serves the Reasoner over an OpenAI-compatible API for image and video understanding.
Start a Nano Reasoner server (publishes the API on port 8000; the first run
downloads the model into ~/.cache/nim):
export NGC_API_KEY=<your_key>
docker run --runtime=nvidia --gpus all \
--shm-size=32GB \
-e NGC_API_KEY="$NGC_API_KEY" \
-e NIM_MODEL_SIZE=nano \
-v ~/.cache/nim:/opt/nim/.cache \
-u $(id -u) \
-p 8000:8000 \
nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0For Cosmos3-Super-Reasoner (the larger model), set -e NIM_MODEL_SIZE=super.
The container serves nvidia/cosmos3-nano-reasoner (or
nvidia/cosmos3-super-reasoner); pass that exact name as the request model, or
resolve it dynamically with client.models.list().
A prebuilt container that serves Cosmos3-Generator Text-to-Video and Image-to-Video
only through POST /v1/infer. The NIM infers mode from the request fields:
non-empty prompt with no image means Text-to-Video; image provided means Image-to-Video. The
response is JSON with a base64-encoded MP4 in b64_video.
It does not expose text-to-image, video-to-video, sound/audio generation, action modes, or transfer controls. Use vLLM-Omni or Cosmos Framework for those broader Generator workflows.
Start a Nano Generator server (default NIM_MODEL_SIZE=nano, NIM_PRECISION=fp8,
NIM_PERF_PROFILE=latency):
export NGC_API_KEY=<your_key>
export LOCAL_NIM_CACHE="${LOCAL_NIM_CACHE:-$HOME/.cache/nim}"
mkdir -p "$LOCAL_NIM_CACHE"
chmod -R 777 "$LOCAL_NIM_CACHE" 2>/dev/null || true
docker run --runtime=nvidia --gpus all \
--shm-size=32GB \
--ulimit nofile=65536:65536 \
-e NGC_API_KEY="$NGC_API_KEY" \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/nvidia/cosmos3-generator:1.0.0For Cosmos3-Super Generator, add -e NIM_MODEL_SIZE=super. Other selection
knobs:
| Env var | Values | Default | Use |
|---|---|---|---|
NIM_MODEL_SIZE |
nano, super |
nano |
Selects 8B Nano or 32B Super |
NIM_PRECISION |
bf16, fp8, nvfp4 |
fp8 |
Selects precision; nvfp4 requires Blackwell |
NIM_PERF_PROFILE |
latency, throughput |
latency |
Optimizes profile selection objective |
NIM_TAGS_SELECTOR |
comma-separated key=value filters |
unset | Advanced profile pinning, e.g. model_size=super,nim_tp=2 |
A quick T2V smoke test:
curl -sS -X POST http://127.0.0.1:8000/v1/infer \
-H 'Accept: application/json' \
-H 'Content-Type: application/json' \
-d '{
"prompt": "A humanoid robot walks through a futuristic warehouse, inspecting shelves of mechanical components.",
"seed": 42,
"guidance_scale": 6.0,
"steps": 35,
"resolution": "256",
"num_output_frames": 25,
"fps": 24.0
}' | jq -r '.b64_video' | base64 -d > /tmp/cosmos3_generator_nim_t2v.mp4For the Cosmos Framework / Diffusers / vLLM venvs, check that PyTorch sees the GPU:
.venv/bin/python - <<'PY'
import torch
print("torch:", torch.__version__)
print("torch cuda:", torch.version.cuda)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
if torch.cuda.is_available():
print("device 0:", torch.cuda.get_device_name(0))
PYFor a vLLM / vLLM-Omni / NIM server, confirm it is serving the model (use the host
port you set with COSMOS3_HOST_PORT or VLLM_PORT):
curl http://localhost:8000/v1/modelsA NIM server also exposes a readiness endpoint that returns 200 once the model
is loaded:
curl http://localhost:8000/v1/health/ready