GPU clusters are registered with the NVCF control plane and managed by the NVCA
Operator. Use the compute-plane Makefile in
deploy/stacks/nvcf-compute-plane for this workflow. It registers one GPU
cluster at a time with nvcf-cli, then installs the operator with the cluster
identity returned by registration. The operator reads that identity from a
local ConfigMap and authenticates through the local OpenBao (Vault) instance.
Clone the public repository and run the compute-plane commands from its root:
git clone https://github.com/nvidia/nvcf.git
cd nvcfBefore installing the NVCA Operator, ensure the following prerequisites are met:
-
The control plane is installed and all core services are running.
-
The NVIDIA GPU Operator is installed on the GPU cluster. The GPU Operator manages the NVIDIA drivers, device plugin, and GPU feature discovery required for workload scheduling. For development or testing environments without physical GPUs, see fake-gpu-operator.
-
(Optional) Install KAI Scheduler for GPU bin-packing and queues. KAI is also the scheduling foundation for gang scheduling and topology-aware scheduling with Grove and Dynamo.
-
GPU Workload Componentsmust be available in a user-managed registry that your Kubernetes cluster can access. SeeGPU Workload Componentsunder self-hosted-artifact-manifest for necessary artifacts and self-hosted-image-mirroring for mirroring instructions. -
nvcf-cliis available on the deployment machine. The compute-plane Makefile calls it during cluster registration. See self-hosted-cli for CLI installation and configuration. -
The SMB CSI driver (
smb.csi.k8s.io) must be installed on the GPU cluster. It is required for NVCA shared model cache storage (samba sidecar). Install it with:helm repo add csi-driver-smb \ https://raw.githubusercontent.com/kubernetes-csi/csi-driver-smb/master/charts helm install csi-driver-smb csi-driver-smb/csi-driver-smb \ -n kube-system --version v1.17.0
In self-managed mode, the cluster identity comes from the CLI registration step, and the operator consumes it:
- You configure the compute-plane Helmfile environment with artifact registry settings and deployment-specific overrides.
- You export the control-plane profile. The profile is the canonical handoff for the control-plane identity, reachable endpoints, host overrides, and transport trust.
- From the repository root, you run
make -C deploy/stacks/nvcf-compute-plane register-cluster(see Register the cluster). The target runsnvcf-cli self-hosted compute-plane registerwith the profile. The CLI discovers the GPU cluster's OIDC issuer and public JWKS, records them with the control plane (ICMS), and writes theclusterID,clusterGroupID, identity source, endpoints, host overrides, and transport trust to a registration values file. - From the repository root, you run
make -C deploy/stacks/nvcf-compute-plane install. Helmfile installshelm-nvca-operatorwith the compute-plane environment and registration values for that cluster, using the same selected kubeconfig context as registration whenCOMPUTE_KUBE_CONTEXTis set. - The Helm chart renders a local ConfigMap (
nvcfbackend-self-managed) holding the cluster identity and the SIS, ReVal, and NATS endpoints (plus any host-header overrides). The operator reads that ConfigMap and creates the NVCFBackend resource. - The operator creates the NVCA agent pod. The agent authenticates to the control plane with a projected service account token (PSAT), which the control plane validates against the issuer and JWKS registered in step 3, and begins managing GPU workloads.
- Runtime secrets are injected by the OpenBao vault-agent sidecar, which authenticates using Kubernetes service account JWT tokens against the local OpenBao instance. Kubernetes image pull secrets are configured separately in the compute-plane environment and namespaces.
sequenceDiagram
actor Operator
participant CP as Control-plane stack
participant Profile as Control-plane profile
participant CLI as nvcf-cli
participant Compute as Compute cluster
participant Values as Registration values
participant Helmfile
Operator->>CLI: profile export
CLI->>CP: read selected environment
CLI->>Profile: write generated profile
Operator->>CLI: init with selected config path
Operator->>CLI: compute-plane register
CLI->>Profile: read endpoints and trust
CLI->>Compute: discover OIDC issuer and JWKS
CLI->>CP: register cluster identity
CLI->>Values: write generated NVCA values
Operator->>Helmfile: make install
Helmfile->>Values: consume generated values
Helmfile->>Compute: install NVCA operator
Create the environment file from the repository root. It provides the chart registry, image registry, image pull secret name, and deployment-specific overrides used by the NVCA operator. The generated registration values provide the control-plane service URLs and host overrides.
cd path/to/nvcf
touch deploy/stacks/nvcf-compute-plane/environments/<environment-name>.yamlglobal:
helm:
sources:
registry: <your-chart-registry>
repository: <your-chart-repository>
image:
registry: <your-image-registry>
repository: <your-image-repository>
imagePullSecrets:
- name: nvcr-pull-secretThe profile exporter derives the service URLs and route-matching host overrides
from the selected control-plane environment, and registration writes them as
the default values for the compute-plane install. Non-empty
global.nvcaOperator.selfManaged.icmsServiceURL,
icmsServiceHostHeaderOverride, revalServiceURL,
revalServiceHostHeaderOverride, natsURL, and natsHostOverride fields in
the selected compute-plane environment take precedence. Verify the effective
compute-reachable endpoints from those inputs resolve from the GPU cluster. See
gateway-routing for service DNS and TLS guidance.
Register the GPU cluster with the control plane before installing the operator.
The nvcf-cli discovers the cluster's OIDC issuer and JWKS and records them
with the control plane, then returns the Helm values the operator needs. See
self-hosted-cli for CLI installation and configuration, and the
Cluster Registration reference for full flag
and output details.
Export the profile from the selected control-plane Helmfile environment. The
default output is
deploy/stacks/self-managed/out/control-plane-profile.yaml:
nvcf-cli self-hosted \
--control-plane-stack deploy/stacks/self-managed \
--env <environment-name> \
control-plane profile export
nvcf-cli --config <path-to-cli-config> initWhen the LLM add-on is disabled, the generated profile and registration values omit its optional router and PKI configuration. Enabling LLM requires the control-plane environment to provide the corresponding managed PKI settings.
Run registration against the GPU cluster kubeconfig. This is required for
multi-cluster installs because the CLI discovers the OIDC issuer and JWKS from
the target Kubernetes cluster. Set COMPUTE_KUBE_CONTEXT to select the GPU
cluster explicitly when the kubeconfig contains multiple contexts.
make -C deploy/stacks/nvcf-compute-plane register-cluster \
CLUSTER_NAME=<gpu-cluster-name> \
CLUSTER_REGION=<region> \
CONTROL_PLANE_PROFILE="$(pwd)/deploy/stacks/self-managed/out/control-plane-profile.yaml" \
KUBECONFIG_FILE=<absolute-path-to-gpu-cluster-kubeconfig> \
COMPUTE_KUBE_CONTEXT=<gpu-cluster-context> \
NVCF_CLI=<absolute-path-to-nvcf-cli> \
NVCF_CLI_CONFIG=<absolute-path-to-cli-config>NVCF_CLI_CONFIG is optional when the default CLI config is correct. The Make
target forwards a configured path to registration but does not run init.
The target writes
registration/<gpu-cluster-name>-register-values.yaml. It carries the cluster
identity and endpoints:
clusterID: <uuid>
clusterGroupID: <uuid>
ncaID: <nca-id>
selfManaged:
region: <region>
identitySource: psat
icmsServiceURL: "http://<GATEWAY_ADDR>"
revalServiceURL: "http://<GATEWAY_ADDR>"
natsURL: "nats://<GATEWAY_ADDR>:4222"selfManaged.identitySource is retained for CLI teardown and is not consumed
by the NVCA Operator chart.
The template, install, and apply targets copy this file into out/ before
running Helmfile.
| Chart | helm-nvca-operator |
|---|---|
| Version | 1.28.0 |
| Namespace | nvca-operator |
| Depends on | All control-plane services and gateway must be running |
The compute-plane Helmfile passes these values to the chart:
| Value | Source |
|---|---|
| Cluster identity | registration/<gpu-cluster-name>-register-values.yaml |
| Operator, agent, image credential helper, and shared storage image repositories | global.image.* in environments/<environment-name>.yaml |
| Control-plane service URLs and host-header overrides | Generated registration values by default; non-empty global.nvcaOperator.selfManaged.* endpoint fields in environments/<environment-name>.yaml take precedence |
| Image pull secrets | global.imagePullSecrets |
Use global.nvcaOperator.nodeSelector, global.nvcaOperator.tolerations, and
global.nvcaOperator.agent.* in the compute-plane environment when you need to
place the operator, agent, or workloads on specific nodes.
The NVCA operator and agent use file watchers for ConfigMap and Secret reconciliation.
Some node images set fs.inotify.max_user_instances to 128, which can be too low for
nodes running the full NVCF stack. When the node exhausts inotify instances, NVCA logs
errors such as failed to create watcher and too many open files. Function creation
or deployment requests that wait for NVCA reconciliation can then return HTTP 500 or
time out.
Before installing the operator, set higher inotify limits on every node. The following DaemonSet sets the values on current nodes and on nodes added later:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-inotify-tuner
namespace: kube-system
spec:
selector:
matchLabels:
app.kubernetes.io/name: node-inotify-tuner
template:
metadata:
labels:
app.kubernetes.io/name: node-inotify-tuner
spec:
hostPID: true
tolerations:
- operator: Exists
priorityClassName: system-node-critical
terminationGracePeriodSeconds: 1
initContainers:
- name: set-sysctl
image: busybox:1.36
securityContext:
privileged: true
command:
- sh
- -c
- |
set -e
mkdir -p /host/etc/sysctl.d
{
echo 'fs.inotify.max_user_instances=8192'
echo 'fs.inotify.max_user_watches=524288'
} > /host/etc/sysctl.d/99-inotify.conf
nsenter -t 1 -m -u -i -n -p -- sysctl -w fs.inotify.max_user_instances=8192
nsenter -t 1 -m -u -i -n -p -- sysctl -w fs.inotify.max_user_watches=524288
volumeMounts:
- name: host
mountPath: /host
containers:
- name: pause
image: registry.k8s.io/pause:3.9
resources:
requests:
cpu: "1m"
memory: "1Mi"
limits:
cpu: "10m"
memory: "16Mi"
volumes:
- name: host
hostPath:
path: /Apply the DaemonSet and wait for it to run on every node:
kubectl apply -f inotify-tuner.yaml
kubectl -n kube-system rollout status ds/node-inotify-tuner --timeout=5mIf your cluster cannot pull public images, mirror busybox:1.36 and
registry.k8s.io/pause:3.9 to your registry and update the image fields before applying
the DaemonSet.
The NVCA operator, NVCA agent, samba sidecar, and image-credential-helper all pull container images from the registry configured in your compute-plane environment file. If that registry is private, create the Kubernetes pull secret before installing the operator.
Use a pre-existing pull secret through `global.imagePullSecrets`. Do not use `generateImagePullSecret`; it does not work in self-managed mode.Create the secret in the GPU cluster namespaces used by the operator and its managed resources:
for ns in nvca-operator nvca-system nvcf-backend; do
kubectl --kubeconfig <gpu-cluster-kubeconfig> \
create namespace "$ns" --dry-run=client -o yaml | kubectl --kubeconfig <gpu-cluster-kubeconfig> apply -f -
kubectl --kubeconfig <gpu-cluster-kubeconfig> \
create secret docker-registry nvcr-pull-secret \
--docker-server=${REGISTRY} \
--docker-username='$oauthtoken' \
--docker-password="$REGISTRY_PASSWORD" \
--namespace="$ns" \
--dry-run=client -o yaml | kubectl --kubeconfig <gpu-cluster-kubeconfig> apply -f -
doneReplace ${REGISTRY} with your container registry (e.g., nvcr.io). For non-NGC
registries, replace --docker-username and --docker-password with your registry
credentials. For NGC (nvcr.io), $REGISTRY_PASSWORD is your NGC Personal Key or
API Key.
Reference the secret in your compute-plane environment:
global:
imagePullSecrets:
- name: nvcr-pull-secretHelmfile passes this value to the operator chart. The operator adds the pull
secret reference to the pods it manages. Pre-creating the secret in
nvca-system and nvcf-backend prevents startup races for operator-managed
resources that pull private images.
Install with the compute-plane Makefile. The command copies
registration/<gpu-cluster-name>-register-values.yaml into out/, then runs
Helmfile against the GPU cluster kubeconfig:
make -C deploy/stacks/nvcf-compute-plane install \
CLUSTER_NAME=<gpu-cluster-name> \
HELMFILE_ENV=<environment-name> \
NCA_ID=<nca-id> \
KUBECONFIG_FILE=<absolute-path-to-gpu-cluster-kubeconfig> \
COMPUTE_KUBE_CONTEXT=<gpu-cluster-context>During installation, Helmfile will:
- Create the operator deployment with vault agent annotations for OpenBao auth.
- Render the
nvcfbackend-self-managedConfigMap from the register values (cluster identity, SIS/ReVal/NATS endpoints, and any host-header overrides). - Start the operator, which reads the ConfigMap and creates the NVCFBackend and NVCA agent deployment.
Check the operator pod is running. The pod runs the operator, the nvca-mirror sidecar, and
the OpenBao vault-agent sidecar (and a cluster-validator init container when network
validation is enabled):
kubectl get pods -n nvca-operator
# Expected:
# NAME READY STATUS RESTARTS AGE
# nvca-operator-... Running 0 1mCheck the cluster identity ConfigMap was rendered from the register values:
kubectl get cm nvcfbackend-self-managed -n nvca-operator \
-o jsonpath='{.data.cluster-dto\.yaml}'
# Expected: cluster-dto.yaml with non-empty clusterId and clusterGroupIdCheck the NVCFBackend resource was created:
kubectl get nvcfbackends -n nvca-operator
# Expected: one NVCFBackend resource with version and health statusCheck the NVCA agent pod is running (the operator creates this automatically):
kubectl get pods -n nvca-system
# Expected:
# NAME READY STATUS RESTARTS AGE
# nvca-... 3/3 Running 0 2mVerify GPU discovery:
kubectl get nvcfbackends -n nvca-operator -o jsonpath='{.items[0].status}' | python3 -m json.tool
# Look for GPU information in the status outputAfter NVCA is healthy, you can install NVIDIA Nsight Operator on the GPU cluster to collect NVIDIA Nsight Systems reports from function pods. See Nsight Profiling for external S3 storage setup, namespace labeling, Kyverno automation, and capture commands.
Use this check only when worker pods on the GPU cluster can reach the
control-plane worker endpoints used by NVCF. For multi-cluster EKS installs,
first validate the GPU cluster by checking nvca-operator rollout and
NVCFBackend agent health. Function deployment and invocation are valid
acceptance checks only after the worker endpoints are reachable from the GPU
cluster.
- Set up environment variables:
# Get the Gateway address (from Step 1)
export GATEWAY_ADDR=$(kubectl get gateway nvcf-gateway -n envoy-gateway -o jsonpath='{.status.addresses[0].value}')
echo "Gateway Address: $GATEWAY_ADDR"- Generate an admin token:
# Generate an admin API token
export NVCF_TOKEN=$(curl -s -X POST "http://${GATEWAY_ADDR}/v1/admin/keys" \
-H "Host: api-keys.${GATEWAY_ADDR}" \
| grep -o '"value":"[^"]*"' | cut -d'"' -f4)
echo "Token generated: ${NVCF_TOKEN:0:20}..."- Create, deploy, and invoke a test function:
# Create a test function
# Replace <YOUR_REGISTRY>/<YOUR_REPO> with your container registry
# This should match the registry you set in the secrets file
curl -s -X POST "http://${GATEWAY_ADDR}/v2/nvcf/functions" \
-H "Host: api.${GATEWAY_ADDR}" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${NVCF_TOKEN}" \
-d '{
"name": "my-echo-function",
"inferenceUrl": "/echo",
"healthUri": "/health",
"inferencePort": 8000,
"containerImage": "<YOUR_REGISTRY>/<YOUR_REPO>/load_tester_supreme:0.0.8"
}' | jq .
# Extract function and version IDs from the response
export FUNCTION_ID=<function-id-from-response>
export FUNCTION_VERSION_ID=<version-id-from-response>
# Deploy the function
# Adjust instanceType and gpu based on your cluster configuration
# Instance Type Examples: NCP.GPU.A10G_1x, NCP.GPU.H100_1x, NCP.GPU.L40S_1x, etc.
# GPU Examples: A10G, H100, L40S, etc.
curl -s -X POST "http://${GATEWAY_ADDR}/v2/nvcf/deployments/functions/${FUNCTION_ID}/versions/${FUNCTION_VERSION_ID}" \
-H "Host: api.${GATEWAY_ADDR}" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${NVCF_TOKEN}" \
-d '{
"deploymentSpecifications": [
{
"instanceType": "NCP.GPU.A10G_1x",
"backend": "nvcf-default",
"gpu": "A10G",
"maxInstances": 1,
"minInstances": 1
}
]
}' | jq .
# Generate an API key for invocation. See the API scope reference for endpoint-specific scope requirements.
# Set expiration to 1 day from now (required field)
EXPIRES_AT=$(date -u -v+1d '+%Y-%m-%dT%H:%M:%SZ' 2>/dev/null || date -u -d '+1 day' '+%Y-%m-%dT%H:%M:%SZ')
SERVICE_ID="nvidia-cloud-functions-ncp-service-id-aketm"
export API_KEY=$(curl -s -X POST "http://${GATEWAY_ADDR}/v1/keys" \
-H "Host: api-keys.${GATEWAY_ADDR}" \
-H "Content-Type: application/json" \
-H "Key-Issuer-Service: nvcf-api" \
-H "Key-Issuer-Id: ${SERVICE_ID}" \
-H "Key-Owner-Id: test@nvcf-api.local" \
-d '{
"description": "test invocation key",
"expires_at": "'"${EXPIRES_AT}"'",
"authorizations": {
"policies": [{
"aud": "'"${SERVICE_ID}"'",
"auds": ["'"${SERVICE_ID}"'"],
"product": "nv-cloud-functions",
"resources": [
{"id": "*", "type": "account-functions"},
{"id": "*", "type": "authorized-functions"}
],
"scopes": ["invoke_function", "list_functions", "queue_details", "list_functions_details"]
}]
},
"audience_service_ids": ["'"${SERVICE_ID}"'"]
}' | jq -r '.value')
echo "API Key: ${API_KEY:0:20}..."
# Wait for deployment to be ready (list functions to see status), then invoke the function
# Uses wildcard subdomain routing: <function-id>.invocation.<gateway-addr>
curl -s -X POST "http://${GATEWAY_ADDR}/echo" \
-H "Host: ${FUNCTION_ID}.invocation.${GATEWAY_ADDR}" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${API_KEY}" \
-d '{"message": "hello world", "repeats": 1}'For invocation, the Host header uses wildcard subdomain routing: <function-id>.invocation.<gateway-addr>.
The URL path should match the function's inferenceUrl (e.g., /echo).
For full HTTP invocation behavior, streaming, and errors, see
Generic HTTP Function Invocation.
You can also use the NVCF CLI for easier function management:
- Create, deploy, and invoke functions with simple commands
- Create or update registry credentials without manual API calls
See self-hosted-cli for installation and usage instructions.
Registration is performed by the CLI, not by the operator. To re-register (for example, after a failed install, or to refresh the recorded issuer and JWKS), re-run the compute-plane registration target and then re-apply the operator with the refreshed values:
make -C deploy/stacks/nvcf-compute-plane register-cluster \
CLUSTER_NAME=<gpu-cluster-name> \
CLUSTER_REGION=<region> \
CONTROL_PLANE_PROFILE="$(pwd)/deploy/stacks/self-managed/out/control-plane-profile.yaml" \
KUBECONFIG_FILE=<absolute-path-to-gpu-cluster-kubeconfig> \
COMPUTE_KUBE_CONTEXT=<gpu-cluster-context> \
NVCF_CLI=<absolute-path-to-nvcf-cli> \
NVCF_CLI_CONFIG=<absolute-path-to-cli-config>
make -C deploy/stacks/nvcf-compute-plane install \
CLUSTER_NAME=<gpu-cluster-name> \
HELMFILE_ENV=<environment-name> \
NCA_ID=<nca-id> \
KUBECONFIG_FILE=<absolute-path-to-gpu-cluster-kubeconfig> \
COMPUTE_KUBE_CONTEXT=<gpu-cluster-context>The registration target rewrites
registration/<gpu-cluster-name>-register-values.yaml. The install target
copies that file into out/ before running Helmfile. The operator reads the
cluster identity at startup, so re-running the install restarts it with the new
values.
To fully remove the NVCA Operator and all associated resources:
If functions are currently deployed on the cluster (pods in the `nvcf-backend` namespace), undeploy them through the NVCF API or CLI before uninstalling the operator. Attempting to delete NVCA while function pods are running can cause finalizers to block namespace deletion. If you encounter stuck resources, see [Handling Stuck Resources] below.-
Delete the NVCFBackend resource. This triggers operator-managed cleanup of the agent deployment, NVCA system pods, and related resources:
kubectl --kubeconfig <gpu-cluster-kubeconfig> \ delete nvcfbackends --all -n nvca-operator --timeout=60s
-
Verify the agent namespace is clean before proceeding:
kubectl --kubeconfig <gpu-cluster-kubeconfig> get pods -n nvca-system # Expected: "No resources found in nvca-system namespace."
-
Destroy the compute-plane Helmfile release:
The cluster identity (`clusterID` and `clusterGroupID`) persists in the control plane (ICMS) and in your registration values file. A reinstall reuses it by re-applying that file. Re-run the supported `register-cluster` target to refresh the generated values before reinstalling.make destroy \ CLUSTER_NAME=<gpu-cluster-name> \ HELMFILE_ENV=<environment-name> \ NCA_ID=<nca-id> \ KUBECONFIG_FILE=<gpu-cluster-kubeconfig>
-
Delete CRDs. This removes all NVCFBackend, MiniService, and StorageRequest custom resources cluster-wide:
kubectl --kubeconfig <gpu-cluster-kubeconfig> delete crd \ nvcfbackends.nvcf.nvidia.io \ miniservices.nvca.nvcf.nvidia.io \ storagerequests.nvca.nvcf.nvidia.io \ --ignore-not-found
-
Delete namespaces:
kubectl --kubeconfig <gpu-cluster-kubeconfig> delete namespace \ nvca-operator nvca-system nvcf-backend nvca-modelcache-init \ --ignore-not-found
If step 1 times out and namespaces remain stuck in Terminating state, or function pods in
nvcf-backend prevent cleanup, use the force-cleanup-script. This script removes
finalizers on stuck NVCA resources, force-deletes function pods, and cleans up all NVCA
namespaces.
# Preview what will be deleted
./force-cleanup-nvcf.sh --dry-run
# Execute the cleanup
./force-cleanup-nvcf.shFor the full script, download link, and detailed usage instructions, see the NVCA Force Cleanup Script appendix in the self-hosted troubleshooting guide.
-
Cluster IDs empty in the ConfigMap after install: The registration values were not applied. Confirm
registration/<gpu-cluster-name>-register-values.yamlexists and that itsclusterIDandclusterGroupIDare populated, then re-runmake install:kubectl get cm nvcfbackend-self-managed -n nvca-operator \ -o jsonpath='{.data.cluster-dto\.yaml}' -
Operator pod not starting: Check the operator logs:
kubectl logs -n nvca-operator -l app.kubernetes.io/name=nvca-operator -c nvca-operator --tail=100
-
Operator or agent logs show
failed to create watcherwithtoo many open files: Increase the node inotify limits with thenode-inotify-tunerDaemonSet in Node inotify limits, then restart the affected pod. -
NVCA agent pod not created: The operator creates the agent pod via the NVCFBackend resource. Check the operator logs for reconciliation errors:
kubectl describe nvcfbackends -n nvca-operator
-
Agent fails to register with SIS (HTTP 401): The control plane could not validate the agent's PSAT against the recorded issuer and JWKS. Re-run
make register-clusterfor the cluster (see Re-registering a cluster) so ICMS has the current issuer and JWKS. Also verify the vault agent sidecar on the agent pod is running and rendering the secrets file:kubectl logs -n nvca-system -l app.kubernetes.io/name=nvca -c vault-agent --tail=10
-
Vault agent sidecar failing: The agent pod needs to authenticate with OpenBao. Verify the vault system is healthy:
kubectl exec -n vault-system openbao-server-0 -- bao status -
No GPUs discovered: Ensure the GPU Operator is installed and GPU nodes have the
nvidia.com/gpuresource advertised:kubectl get nodes -o custom-columns="NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu"