You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
NVSwitch health events reach the datastore and no Prometheus metric, by two independent routes. #1825 fixed the collection side; this is about what happens to the events afterwards.
Up front, the limitation in this report: I cannot demonstrate the runtime symptom. Our fleet is GB200 NVL72, where dcgmi discovery -l reports 0 NvSwitches found on every compute node, so the NVSwitch branch never executes for us and we have zero NVSwitch events to observe. Everything below is from reading the code on main, and the inference about runtime behaviour is labelled as such. Someone on an HGX or DGX baseboard with on-board NVSwitch can confirm it in a minute; I cannot.
1. STORE_ONLY is hardcoded at all three NVSwitch emit sites
health-monitors/gpu-health-monitor/gpu_health_monitor/platform_connector/platform_connector.py, lines 632, 658 and 683, each paired with componentClass="NVSWITCH" at 622, 647 and 676:
commented "NVSwitch remediation is deferred until downstream handling is safe".
STORE_ONLY events are filtered before the Prometheus connector counts them, so they never appear in health_events_total. That much is a deliberate choice about remediation, and I am not arguing with deferring remediation. The side effect is that it also removes them from the counter that alerting is built on, which is a different concern from whether anything should act on them.
Comparable checks expose this as a value. gpuPowerBrakeStoreOnly lets an operator choose observe-only versus acted-on for the power brake. The NVSwitch path has no equivalent knob.
2. The NVSwitch branch never populates dcgm_health_active_events
This is the one I would fix first, because it looks like an oversight rather than a decision, and because it changes no remediation semantics at all.
That gauge is written from pending_metric_updates, drained at line 699:
Every append to that list is in the GPU branch, at lines 501, 535 and 588. The NVSwitch branch (roughly 600 to 690) appends to health_events and to pending_cache_updates, and never to pending_metric_updates.
So the one metric that survives STORE_ONLY has no NVSwitch series either.
Combined effect
Inference, not observation: on a system where DCGM does enumerate switches, a DCGM_FR_NVSWITCH_FATAL_ERROR is detected, cached, and written to MongoDB, and is then invisible to health_events_total (filtered as STORE_ONLY) and to dcgm_health_active_events (never recorded). No alert is possible from either metric, and an operator watching a dashboard sees nothing.
That the fault is correctly attributed after #1825 makes this more noticeable, not less: the data is now right and still cannot be seen.
Make the NVSwitch STORE_ONLY configurable, in the shape of gpuPowerBrakeStoreOnly. Detection-only deployments want these events counted while remediation stays off. I have lower confidence this is wanted, given the comment reads as deliberate, which is why it is second and separable.
Happy to write (1) if it is welcome, with the caveat that I cannot test it on our hardware and would be relying on CI plus a reviewer with switches.
NVSentinel v1.24.0, chart v1.24.0, MongoDB datastore, detection-only. Line numbers verified against origin/main at the time of filing and are identical to the v1.24.0 tag.
Description
NVSwitch health events reach the datastore and no Prometheus metric, by two independent routes. #1825 fixed the collection side; this is about what happens to the events afterwards.
Up front, the limitation in this report: I cannot demonstrate the runtime symptom. Our fleet is GB200 NVL72, where
dcgmi discovery -lreports0 NvSwitches foundon every compute node, so the NVSwitch branch never executes for us and we have zero NVSwitch events to observe. Everything below is from reading the code onmain, and the inference about runtime behaviour is labelled as such. Someone on an HGX or DGX baseboard with on-board NVSwitch can confirm it in a minute; I cannot.1.
STORE_ONLYis hardcoded at all three NVSwitch emit siteshealth-monitors/gpu-health-monitor/gpu_health_monitor/platform_connector/platform_connector.py, lines 632, 658 and 683, each paired withcomponentClass="NVSWITCH"at 622, 647 and 676:commented "NVSwitch remediation is deferred until downstream handling is safe".
STORE_ONLYevents are filtered before the Prometheus connector counts them, so they never appear inhealth_events_total. That much is a deliberate choice about remediation, and I am not arguing with deferring remediation. The side effect is that it also removes them from the counter that alerting is built on, which is a different concern from whether anything should act on them.Comparable checks expose this as a value.
gpuPowerBrakeStoreOnlylets an operator choose observe-only versus acted-on for the power brake. The NVSwitch path has no equivalent knob.2. The NVSwitch branch never populates
dcgm_health_active_eventsThis is the one I would fix first, because it looks like an oversight rather than a decision, and because it changes no remediation semantics at all.
That gauge is written from
pending_metric_updates, drained at line 699:Every append to that list is in the GPU branch, at lines 501, 535 and 588. The NVSwitch branch (roughly 600 to 690) appends to
health_eventsand topending_cache_updates, and never topending_metric_updates.So the one metric that survives
STORE_ONLYhas no NVSwitch series either.Combined effect
Inference, not observation: on a system where DCGM does enumerate switches, a
DCGM_FR_NVSWITCH_FATAL_ERRORis detected, cached, and written to MongoDB, and is then invisible tohealth_events_total(filtered as STORE_ONLY) and todcgm_health_active_events(never recorded). No alert is possible from either metric, and an operator watching a dashboard sees nothing.That the fault is correctly attributed after #1825 makes this more noticeable, not less: the data is now right and still cannot be seen.
Suggested fixes, in the order I would do them
pending_metric_updatesin the NVSwitch branch, mirroring the GPU branch. Two notes from having made the same class of change in feat: label dcgm_health_active_events with the error code #1874: the clear path has to zero the same label set that was set, or a series latches at 1 for the process lifetime; and since feat: label dcgm_health_active_events with the error code #1874 the gauge carrieserror_code, so a check latched by several codes needs each one cleared.STORE_ONLYconfigurable, in the shape ofgpuPowerBrakeStoreOnly. Detection-only deployments want these events counted while remediation stays off. I have lower confidence this is wanted, given the comment reads as deliberate, which is why it is second and separable.Happy to write (1) if it is welcome, with the caveat that I cannot test it on our hardware and would be relying on CI plus a reviewer with switches.
Not duplicates, for the record
Environment
NVSentinel v1.24.0, chart v1.24.0, MongoDB datastore, detection-only. Line numbers verified against
origin/mainat the time of filing and are identical to the v1.24.0 tag.