You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 85ae398
Browse filesBrowse the repository at this point in the historyBrowse files
@@ -301,6 +302,14 @@ These metrics track GPU health events detected via DCGM (Data Center GPU Manager
301
302
302
303
---
303
304
305
+
### NIC Health Monitor
306
+
307
+
| Metric Name | Type | Labels | Description |
308
+
|------------|------|--------|-------------|
309
+
|`nic_health_monitor_poll_cycle_last_completed_timestamp_seconds`| Gauge |`node`, `category`| Unix timestamp initialized at monitor startup and advanced after each completed poll cycle (`state` or `counter`). Use `time() - metric` to measure how long the category has gone without completing a poll, including a stall in the first poll. |
310
+
311
+
---
312
+
304
313
### Syslog Health Monitor
305
314
306
315
The syslog health monitor tracks GPU-related errors detected from system logs.
Copy file name to clipboardExpand all lines: docs/configuration/gpu-health-monitor.md
+6-1Lines changed: 6 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -253,6 +253,7 @@ gpu-health-monitor:
253
253
dcgmFieldsMonitoring:
254
254
gpuTempLimitMonitoringEnabled: true
255
255
gpuTempLimitStoreOnly: true
256
+
gpuTempLimitMinConsecutivePolls: 3
256
257
```
257
258
258
259
### gpuTempLimitMonitoringEnabled
@@ -263,6 +264,10 @@ Enables the watch. On by default.
263
264
264
265
Dry run. When true, this check's events are emitted with `processingStrategy=STORE_ONLY`, so they are persisted and exported as metrics but excluded from the remediation pipeline: no node condition and no cordon. Defaults to true, so the watch is observable before it can act. Set it to `false` once you have confirmed the thresholds suit your hardware and cooling.
265
266
267
+
### gpuTempLimitMinConsecutivePolls
268
+
269
+
Consecutive polls with the margin below the slowdown threshold before the GPU is failed. Defaults to 3. The margin is read as one sample per poll, and under load it can move by tens of degrees within seconds, so a single sample past the threshold is usually a transient that the GPU's own hardware slowdown has already handled. A sustained excursion is the actionable case. A sample at or above the threshold resets the counter, and `1` fails on first observation. A GPU with no usable sample is skipped and keeps its counter, so a gap in DCGM data neither raises nor clears a finding. While the counter is below the threshold the GPU is not reported either way, so a restart cannot publish a healthy event for a GPU that is still past the threshold.
270
+
266
271
To interpret a firing check, see the [GPU Thermal Margin runbook](../runbooks/gpu-thermal-margin.md).
267
272
268
273
## NVLink Suppression on Unbridged PCIe Cards
@@ -304,7 +309,7 @@ Dry run. When true, this check's events are emitted with `processingStrategy=STO
304
309
305
310
### gpuPowerBrakeMinConsecutivePolls
306
311
307
-
Consecutive polls with the bit set before the GPU is failed. A brake asserted for a single poll can be a load transient; a sustained assertion is the actionable case. A clear resets the counter, so a flapping brake never accumulates to a failure. `1` fails on first observation. A GPU with no usable sample is skipped and keeps its counter, so a gap in DCGM data neither raises nor clears a finding. This includes DCGM's int64 "no data" sentinels, whose low byte has the brake bit set and which would otherwise read as an assertion.
312
+
Consecutive polls with the bit set before the GPU is failed. A brake asserted for a single poll can be a load transient; a sustained assertion is the actionable case. A clear resets the counter, so a flapping brake never accumulates to a failure. `1` fails on first observation. A GPU with no usable sample is skipped and keeps its counter, so a gap in DCGM data neither raises nor clears a finding. This includes DCGM's int64 "no data" sentinels, whose low byte has the brake bit set and which would otherwise read as an assertion. While the counter is below the threshold the GPU is not reported either way, so a restart cannot publish a healthy event for a GPU whose brake is still asserted.
This configuration allows a client to create a ValidationRequest that runs dcgm-diag-test as a Kubernetes Job via the k8s-job-provider. Before the test group starts, the targeted node must report allocatable GPU capacity and not be under quarantine, as enforced by readinessCriteria. If the ValidationRequest does not specify spec.tests, defaultTests provides which tests to run in the request.
76
+
This configuration allows a client to create a ValidationRequest that runs dcgm-diag-test as a Kubernetes Job via the k8s-job-provider. Before the test group starts, the targeted node must report allocatable GPU capacity and not be under quarantine, as enforced by readinessCriteria. The gpu-allocatable criterion accepts either source of GPU capacity: the device plugin, which sets `nvidia.com/gpu` in the node allocatable resources, or the `gpu.nvidia.com` DRA driver, which publishes the GPUs in ResourceSlices in GPU Operator GPUCluster mode. If the ValidationRequest does not specify spec.tests, defaultTests provides which tests to run in the request.
68
77
69
78
## Enabling New Node Validation
70
79
@@ -281,7 +290,7 @@ status:
281
290
| Key | Type | Purpose |
282
291
|---|---|---|
283
292
| defaultTests | []string | The default set of tests run against ValidationRequests which do not include any tests |
284
-
| readinessCriteria | []CriteriaSpec | A set of CEL expressions which must all evaluate to true before a validation test can be started on a given node. Each entry is evaluated against an environment containing the node being validated. If an operator is externally applying a node cordon or taint and would like to block validation until these are applied, they can add these properties to readinessCriteria. If not met, this blocks a node from starting validation and fails validation if the criteria were initially met and then reverted |
293
+
| readinessCriteria | []CriteriaSpec | A set of CEL expressions which must all evaluate to true before a validation test can be started on a given node. Each entry is evaluated against an environment containing the node being validated (`node`) and the ResourceSlices whose `spec.nodeName` is that node (`resourceSlices`). When a criterion reads `resourceSlices`, the lifecycle-manager watches ResourceSlices so a slice published after the last node update still unblocks a pending request. It derives the drivers to watch from the expressions: slice events are filtered to the `spec.driver` values compared with string literals, as `s.spec.driver == "gpu.nvidia.com"` or `s.spec.driver in [...]`. A criterion that reads `spec.driver` any other way, or reads slices without testing the driver, turns the filter off and every driver's slice events are processed (logged at startup). Filtering only affects which events wake the controller; the `resourceSlices` variable always holds all of the node's slices. If an operator is externally applying a node cordon or taint and would like to block validation until these are applied, they can add these properties to readinessCriteria. If not met, this blocks a node from starting validation and fails validation if the criteria were initially met and then reverted |
285
294
| maxConcurrentGroups | int | The maximum number of test groups that may run concurrently. Groups are additionally constrained by node overlap. Two groups that share a node never run at the same time regardless of this setting |
286
295
| templateMountPath | string | The directory from which templateFile paths are resolved |
287
296
| providers | map[string]ProviderConfig | Test provider settings that apply to all tests using this provider, keyed by the name tests[].provider references (see below) |
@@ -321,7 +330,7 @@ status:
321
330
| Key | Type | Purpose |
322
331
|---|---|---|
323
332
| condition | string | The name of the node condition the controller uses to track whether a node has already been validated. For a node to be targeted, this condition must be absent or false and every criteria expression must evaluate to true. Once a ValidationRequest is created, the controller sets this condition to True on the node so that subsequent evaluations no longer match |
324
-
| criteria | []CriteriaSpec | A set of CEL expressions evaluated against each node to determine whether it requires new node validation. All expressions must evaluate to true, along with the condition check above. The CEL environment exposes the node being validated |
333
+
| criteria | []CriteriaSpec | A set of CEL expressions evaluated against each node to determine whether it requires new node validation. All expressions must evaluate to true, along with the condition check above. The CEL environment exposes the node being validated (`node`) and its node-local ResourceSlices (`resourceSlices`), the same as readinessCriteria |
325
334
| newNodeTests | []string | The list of tests to run for new nodes. These take precedence over defaultTests when a ValidationRequest is created for a new node |
326
335
| batchPeriodSeconds | int64 | The window during which the controller collects eligible new nodes before creating ValidationRequests for them as a batch. Only applies to new node validation |
Copy file name to clipboardExpand all lines: docs/runbooks/gpu-thermal-margin.md
+2-1Lines changed: 2 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -2,7 +2,7 @@
2
2
3
3
## Overview
4
4
5
-
`GpuThermalMarginWatch` monitors each GPU's live thermal margin (DCGM field 153, `DCGM_FI_DEV_GPU_TEMP_LIMIT`) against the per-SKU hardware-slowdown T.Limit offset published by the metadata-collector (NVML field 194, `FI_DEV_TEMPERATURE_SLOWDOWN_TLIMIT`). When a GPU's margin drops below its slowdown threshold, the GPU is at or past the temperature at which the hardware engages thermal slowdown, the fatal event for GpuThermalMarginWatch occurs. The feature is described in [ADR-042: GPU Thermal Margin](../designs/042-gpu-temp-limit-field-monitoring.md).
5
+
`GpuThermalMarginWatch` monitors each GPU's live thermal margin (DCGM field 153, `DCGM_FI_DEV_GPU_TEMP_LIMIT`) against the per-SKU hardware-slowdown T.Limit offset published by the metadata-collector (NVML field 194, `FI_DEV_TEMPERATURE_SLOWDOWN_TLIMIT`). When a GPU's margin stays below its slowdown threshold for `gpuTempLimitMinConsecutivePolls` consecutive polls (3 by default), the GPU has been at or past the temperature at which the hardware engages thermal slowdown, and the fatal event for GpuThermalMarginWatch occurs. The feature is described in [ADR-042: GPU Thermal Margin](../designs/042-gpu-temp-limit-field-monitoring.md).
6
6
7
7
**Key points:**
8
8
@@ -124,5 +124,6 @@ The condition clears automatically once the live margin returns to at or above t
124
124
The check is configured through the `gpu-health-monitor` Helm chart, which renders the `[dcgmfieldsmonitoring]` section of `config.ini`:
125
125
126
126
- Enable the check: Helm value `dcgmFieldsMonitoring.gpuTempLimitMonitoringEnabled` renders `gputemplimitmonitoringenabled`.
127
+
- Consecutive polls required to fail: Helm value `dcgmFieldsMonitoring.gpuTempLimitMinConsecutivePolls` renders `gputemplimitminconsecutivepolls`.
127
128
128
129
The per-GPU threshold (`slowdown_tlimit_c`) is not a Helm value. It is collected at runtime by the metadata-collector and written to `/var/lib/nvsentinel/gpu_metadata.json`.
0 commit comments