Skip to content

Commit 85ae398

Browse files
authored
Merge branch 'main' into fix/fq-retain-taints-during-validation
2 parents 3c30a19 + 2ef568a commit 85ae398

23 files changed

Lines changed: 921 additions & 85 deletions

File tree

‎distros/kubernetes/nvsentinel/charts/gpu-health-monitor/templates/configmap.yaml‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -37,6 +37,7 @@ data:
3737
[dcgmfieldsmonitoring]
3838
gputemplimitmonitoringenabled = {{ .Values.dcgmFieldsMonitoring.gpuTempLimitMonitoringEnabled }}
3939
gputemplimitstoreonly = {{ .Values.dcgmFieldsMonitoring.gpuTempLimitStoreOnly }}
40+
gputemplimitminconsecutivepolls = {{ .Values.dcgmFieldsMonitoring.gpuTempLimitMinConsecutivePolls }}
4041
gpupowerbrakemonitoringenabled = {{ .Values.dcgmFieldsMonitoring.gpuPowerBrakeMonitoringEnabled }}
4142
gpupowerbrakestoreonly = {{ .Values.dcgmFieldsMonitoring.gpuPowerBrakeStoreOnly }}
4243
gpupowerbrakeminconsecutivepolls = {{ .Values.dcgmFieldsMonitoring.gpuPowerBrakeMinConsecutivePolls }}

‎distros/kubernetes/nvsentinel/charts/gpu-health-monitor/values.yaml‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -90,6 +90,11 @@ dcgmFieldsMonitoring:
9090
# Dry-run for GpuThermalMarginWatch. When true, this check's
9191
# events are emitted with processingStrategy=STORE_ONLY
9292
gpuTempLimitStoreOnly: true
93+
# Consecutive polls with the margin below the slowdown T.Limit before the GPU is
94+
# failed. The margin is sampled once per poll and swings by tens of degrees within
95+
# seconds under load, so one sample past the threshold is usually a transient.
96+
# 1 fails on first observation.
97+
gpuTempLimitMinConsecutivePolls: 3
9398
# Enable GpuPowerBrakeWatch: fails a GPU whose clocks-event-reasons mask has the
9499
# HW power brake bit (0x80) set, i.e. the power delivery path is forcing clocks
95100
# down. DCGM's POWER health watch does not report this; its dominant code

‎distros/kubernetes/nvsentinel/charts/lifecycle-manager/templates/clusterrole.yaml‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -55,6 +55,13 @@ rules:
5555
verbs:
5656
- update
5757
- patch
58+
- apiGroups:
59+
- resource.k8s.io
60+
resources:
61+
- resourceslices
62+
verbs:
63+
- list
64+
- watch
5865
- apiGroups:
5966
- coordination.k8s.io
6067
resources:

‎distros/kubernetes/nvsentinel/charts/lifecycle-manager/values.yaml‎

Lines changed: 7 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -63,11 +63,16 @@ config:
6363
# since a toleration for this taint is always added in test provider templates which
6464
# consume the Tolerations field. Setting cordon.remove to true will remove the node
6565
# cordon (and this taint) if present.
66+
# Expressions can read the node being validated (node) and the node's ResourceSlices (resourceSlices).
6667
readinessCriteria:
68+
# The device plugin advertises nvidia.com/gpu in node.status.allocatable. In GPU Operator GPUCluster (DRA)
69+
# mode, there is no device plugin and the gpu.nvidia.com DRA driver publishes the GPUs in ResourceSlices.
6770
- name: gpu-allocatable
6871
expression: >-
69-
has(node.status.allocatable) && "nvidia.com/gpu" in node.status.allocatable &&
70-
quantity(node.status.allocatable["nvidia.com/gpu"]) > 0
72+
(has(node.status.allocatable) && "nvidia.com/gpu" in node.status.allocatable &&
73+
quantity(node.status.allocatable["nvidia.com/gpu"]) > 0) ||
74+
resourceSlices.exists(s, s.spec.driver == "gpu.nvidia.com" && has(s.spec.devices) &&
75+
size(s.spec.devices) > 0)
7176
- name: not-under-quarantine
7277
expression: >-
7378
!(has(node.metadata.annotations) &&

‎docs/METRICS.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -13,6 +13,7 @@ This document outlines all Prometheus metrics exposed by NVSentinel components.
1313
- [Health Monitors](#health-monitors)
1414
- [Health Event Publisher](#health-event-publisher)
1515
- [GPU Health Monitor](#gpu-health-monitor)
16+
- [NIC Health Monitor](#nic-health-monitor)
1617
- [Syslog Health Monitor](#syslog-health-monitor)
1718
- [CSP Health Monitor](#csp-health-monitor)
1819
- [Change Stream Metrics](#change-stream-metrics)
@@ -301,6 +302,14 @@ These metrics track GPU health events detected via DCGM (Data Center GPU Manager
301302

302303
---
303304

305+
### NIC Health Monitor
306+
307+
| Metric Name | Type | Labels | Description |
308+
|------------|------|--------|-------------|
309+
| `nic_health_monitor_poll_cycle_last_completed_timestamp_seconds` | Gauge | `node`, `category` | Unix timestamp initialized at monitor startup and advanced after each completed poll cycle (`state` or `counter`). Use `time() - metric` to measure how long the category has gone without completing a poll, including a stall in the first poll. |
310+
311+
---
312+
304313
### Syslog Health Monitor
305314

306315
The syslog health monitor tracks GPU-related errors detected from system logs.

‎docs/configuration/gpu-health-monitor.md‎

Lines changed: 6 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -253,6 +253,7 @@ gpu-health-monitor:
253253
dcgmFieldsMonitoring:
254254
gpuTempLimitMonitoringEnabled: true
255255
gpuTempLimitStoreOnly: true
256+
gpuTempLimitMinConsecutivePolls: 3
256257
```
257258

258259
### gpuTempLimitMonitoringEnabled
@@ -263,6 +264,10 @@ Enables the watch. On by default.
263264

264265
Dry run. When true, this check's events are emitted with `processingStrategy=STORE_ONLY`, so they are persisted and exported as metrics but excluded from the remediation pipeline: no node condition and no cordon. Defaults to true, so the watch is observable before it can act. Set it to `false` once you have confirmed the thresholds suit your hardware and cooling.
265266

267+
### gpuTempLimitMinConsecutivePolls
268+
269+
Consecutive polls with the margin below the slowdown threshold before the GPU is failed. Defaults to 3. The margin is read as one sample per poll, and under load it can move by tens of degrees within seconds, so a single sample past the threshold is usually a transient that the GPU's own hardware slowdown has already handled. A sustained excursion is the actionable case. A sample at or above the threshold resets the counter, and `1` fails on first observation. A GPU with no usable sample is skipped and keeps its counter, so a gap in DCGM data neither raises nor clears a finding. While the counter is below the threshold the GPU is not reported either way, so a restart cannot publish a healthy event for a GPU that is still past the threshold.
270+
266271
To interpret a firing check, see the [GPU Thermal Margin runbook](../runbooks/gpu-thermal-margin.md).
267272

268273
## NVLink Suppression on Unbridged PCIe Cards
@@ -304,7 +309,7 @@ Dry run. When true, this check's events are emitted with `processingStrategy=STO
304309

305310
### gpuPowerBrakeMinConsecutivePolls
306311

307-
Consecutive polls with the bit set before the GPU is failed. A brake asserted for a single poll can be a load transient; a sustained assertion is the actionable case. A clear resets the counter, so a flapping brake never accumulates to a failure. `1` fails on first observation. A GPU with no usable sample is skipped and keeps its counter, so a gap in DCGM data neither raises nor clears a finding. This includes DCGM's int64 "no data" sentinels, whose low byte has the brake bit set and which would otherwise read as an assertion.
312+
Consecutive polls with the bit set before the GPU is failed. A brake asserted for a single poll can be a load transient; a sustained assertion is the actionable case. A clear resets the counter, so a flapping brake never accumulates to a failure. `1` fails on first observation. A GPU with no usable sample is skipped and keeps its counter, so a gap in DCGM data neither raises nor clears a finding. This includes DCGM's int64 "no data" sentinels, whose low byte has the brake bit set and which would otherwise read as an assertion. While the counter is below the threshold the GPU is not reported either way, so a restart cannot publish a healthy event for a GPU whose brake is still asserted.
308313

309314
## DCGM Startup Gate
310315

‎docs/configuration/validation.md‎

Lines changed: 15 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -27,8 +27,10 @@ lifecycle-manager:
2727
readinessCriteria:
2828
- name: gpu-allocatable
2929
expression: >-
30-
has(node.status.allocatable) && "nvidia.com/gpu" in node.status.allocatable &&
31-
quantity(node.status.allocatable["nvidia.com/gpu"]) > 0
30+
(has(node.status.allocatable) && "nvidia.com/gpu" in node.status.allocatable &&
31+
quantity(node.status.allocatable["nvidia.com/gpu"]) > 0) ||
32+
resourceSlices.exists(s, s.spec.driver == "gpu.nvidia.com" && has(s.spec.devices) &&
33+
size(s.spec.devices) > 0)
3234
- name: not-under-quarantine
3335
expression: >-
3436
!(has(node.metadata.annotations) &&
@@ -58,13 +60,20 @@ lifecycle-manager:
5860
command:
5961
- sh
6062
- -c
61-
- dcgmi diag --host "nvidia-dcgm.gpu-operator.svc:5555" --run 2 --json
63+
- |
64+
# One DCGM Service exists per cluster: nvidia-dcgm-dra in GPU Operator GPUCluster (DRA) mode,
65+
# nvidia-dcgm otherwise. Use the first name that resolves.
66+
for h in nvidia-dcgm-dra.gpu-operator.svc nvidia-dcgm.gpu-operator.svc; do
67+
getent hosts "$h" >/dev/null 2>&1 && exec dcgmi diag --host "$h:5555" --run 2 --json
68+
done
69+
echo "no DCGM hostengine Service resolved" >&2
70+
exit 1
6271
supportsBatchingNodes: false
6372
minimumNodesPerBatch: 1
6473
batchFailurePolicy: fail
6574
```
6675
67-
This configuration allows a client to create a ValidationRequest that runs dcgm-diag-test as a Kubernetes Job via the k8s-job-provider. Before the test group starts, the targeted node must report allocatable GPU capacity and not be under quarantine, as enforced by readinessCriteria. If the ValidationRequest does not specify spec.tests, defaultTests provides which tests to run in the request.
76+
This configuration allows a client to create a ValidationRequest that runs dcgm-diag-test as a Kubernetes Job via the k8s-job-provider. Before the test group starts, the targeted node must report allocatable GPU capacity and not be under quarantine, as enforced by readinessCriteria. The gpu-allocatable criterion accepts either source of GPU capacity: the device plugin, which sets `nvidia.com/gpu` in the node allocatable resources, or the `gpu.nvidia.com` DRA driver, which publishes the GPUs in ResourceSlices in GPU Operator GPUCluster mode. If the ValidationRequest does not specify spec.tests, defaultTests provides which tests to run in the request.
6877

6978
## Enabling New Node Validation
7079

@@ -281,7 +290,7 @@ status:
281290
| Key | Type | Purpose |
282291
|---|---|---|
283292
| defaultTests | []string | The default set of tests run against ValidationRequests which do not include any tests |
284-
| readinessCriteria | []CriteriaSpec | A set of CEL expressions which must all evaluate to true before a validation test can be started on a given node. Each entry is evaluated against an environment containing the node being validated. If an operator is externally applying a node cordon or taint and would like to block validation until these are applied, they can add these properties to readinessCriteria. If not met, this blocks a node from starting validation and fails validation if the criteria were initially met and then reverted |
293+
| readinessCriteria | []CriteriaSpec | A set of CEL expressions which must all evaluate to true before a validation test can be started on a given node. Each entry is evaluated against an environment containing the node being validated (`node`) and the ResourceSlices whose `spec.nodeName` is that node (`resourceSlices`). When a criterion reads `resourceSlices`, the lifecycle-manager watches ResourceSlices so a slice published after the last node update still unblocks a pending request. It derives the drivers to watch from the expressions: slice events are filtered to the `spec.driver` values compared with string literals, as `s.spec.driver == "gpu.nvidia.com"` or `s.spec.driver in [...]`. A criterion that reads `spec.driver` any other way, or reads slices without testing the driver, turns the filter off and every driver's slice events are processed (logged at startup). Filtering only affects which events wake the controller; the `resourceSlices` variable always holds all of the node's slices. If an operator is externally applying a node cordon or taint and would like to block validation until these are applied, they can add these properties to readinessCriteria. If not met, this blocks a node from starting validation and fails validation if the criteria were initially met and then reverted |
285294
| maxConcurrentGroups | int | The maximum number of test groups that may run concurrently. Groups are additionally constrained by node overlap. Two groups that share a node never run at the same time regardless of this setting |
286295
| templateMountPath | string | The directory from which templateFile paths are resolved |
287296
| providers | map[string]ProviderConfig | Test provider settings that apply to all tests using this provider, keyed by the name tests[].provider references (see below) |
@@ -321,7 +330,7 @@ status:
321330
| Key | Type | Purpose |
322331
|---|---|---|
323332
| condition | string | The name of the node condition the controller uses to track whether a node has already been validated. For a node to be targeted, this condition must be absent or false and every criteria expression must evaluate to true. Once a ValidationRequest is created, the controller sets this condition to True on the node so that subsequent evaluations no longer match |
324-
| criteria | []CriteriaSpec | A set of CEL expressions evaluated against each node to determine whether it requires new node validation. All expressions must evaluate to true, along with the condition check above. The CEL environment exposes the node being validated |
333+
| criteria | []CriteriaSpec | A set of CEL expressions evaluated against each node to determine whether it requires new node validation. All expressions must evaluate to true, along with the condition check above. The CEL environment exposes the node being validated (`node`) and its node-local ResourceSlices (`resourceSlices`), the same as readinessCriteria |
325334
| newNodeTests | []string | The list of tests to run for new nodes. These take precedence over defaultTests when a ValidationRequest is created for a new node |
326335
| batchPeriodSeconds | int64 | The window during which the controller collects eligible new nodes before creating ValidationRequests for them as a batch. Only applies to new node validation |
327336

‎docs/runbooks/gpu-thermal-margin.md‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,7 @@
22

33
## Overview
44

5-
`GpuThermalMarginWatch` monitors each GPU's live thermal margin (DCGM field 153, `DCGM_FI_DEV_GPU_TEMP_LIMIT`) against the per-SKU hardware-slowdown T.Limit offset published by the metadata-collector (NVML field 194, `FI_DEV_TEMPERATURE_SLOWDOWN_TLIMIT`). When a GPU's margin drops below its slowdown threshold, the GPU is at or past the temperature at which the hardware engages thermal slowdown, the fatal event for GpuThermalMarginWatch occurs. The feature is described in [ADR-042: GPU Thermal Margin](../designs/042-gpu-temp-limit-field-monitoring.md).
5+
`GpuThermalMarginWatch` monitors each GPU's live thermal margin (DCGM field 153, `DCGM_FI_DEV_GPU_TEMP_LIMIT`) against the per-SKU hardware-slowdown T.Limit offset published by the metadata-collector (NVML field 194, `FI_DEV_TEMPERATURE_SLOWDOWN_TLIMIT`). When a GPU's margin stays below its slowdown threshold for `gpuTempLimitMinConsecutivePolls` consecutive polls (3 by default), the GPU has been at or past the temperature at which the hardware engages thermal slowdown, and the fatal event for GpuThermalMarginWatch occurs. The feature is described in [ADR-042: GPU Thermal Margin](../designs/042-gpu-temp-limit-field-monitoring.md).
66

77
**Key points:**
88

@@ -124,5 +124,6 @@ The condition clears automatically once the live margin returns to at or above t
124124
The check is configured through the `gpu-health-monitor` Helm chart, which renders the `[dcgmfieldsmonitoring]` section of `config.ini`:
125125

126126
- Enable the check: Helm value `dcgmFieldsMonitoring.gpuTempLimitMonitoringEnabled` renders `gputemplimitmonitoringenabled`.
127+
- Consecutive polls required to fail: Helm value `dcgmFieldsMonitoring.gpuTempLimitMinConsecutivePolls` renders `gputemplimitminconsecutivepolls`.
127128

128129
The per-GPU threshold (`slowdown_tlimit_c`) is not a Helm value. It is collected at runtime by the metadata-collector and written to `/var/lib/nvsentinel/gpu_metadata.json`.

‎health-monitors/gpu-health-monitor/gpu_health_monitor/cli.py‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -207,17 +207,22 @@ def cli(
207207

208208
thermal_margin_enabled = False
209209
thermal_margin_store_only = False
210+
thermal_margin_min_consecutive_polls = 1
210211
power_brake_enabled = False
211212
power_brake_store_only = False
212213
power_brake_min_consecutive_polls = 1
213214
if config.has_section("dcgmfieldsmonitoring"):
214215
fields_monitoring_config = config["dcgmfieldsmonitoring"]
215216
thermal_margin_enabled = fields_monitoring_config.getboolean("gputemplimitmonitoringenabled", fallback=False)
216217
thermal_margin_store_only = fields_monitoring_config.getboolean("gputemplimitstoreonly", fallback=False)
218+
thermal_margin_min_consecutive_polls = fields_monitoring_config.getint(
219+
"gputemplimitminconsecutivepolls", fallback=1
220+
)
217221
log.info(
218-
"GpuThermalMarginWatch field monitor: enabled=%s store_only=%s",
222+
"GpuThermalMarginWatch field monitor: enabled=%s store_only=%s min_consecutive_polls=%s",
219223
thermal_margin_enabled,
220224
thermal_margin_store_only,
225+
thermal_margin_min_consecutive_polls,
221226
)
222227

223228
power_brake_enabled = fields_monitoring_config.getboolean("gpupowerbrakemonitoringenabled", fallback=False)
@@ -361,6 +366,7 @@ def process_exit_signal(signum, frame):
361366
poll_interval_seconds=poll_interval,
362367
dcgm_k8s_service_enabled=dcgm_k8s_service_enabled,
363368
thermal_margin_enabled=thermal_margin_enabled,
369+
thermal_margin_min_consecutive_polls=thermal_margin_min_consecutive_polls,
364370
dcgm_mode=dcgm_mode,
365371
suppressed_error_codes=suppressed_error_codes,
366372
suppress_unbridged_pcie_nvlink_down=suppress_nvlink_down_unbridged_pcie,

0 commit comments

Comments
 (0)