Skip to content

[Bug]: NVSwitch health events reach no Prometheus metric: STORE_ONLY hardcoded and dcgm_health_active_events never populated #1902

Description

@lfriedman-netllama

Description

NVSwitch health events reach the datastore and no Prometheus metric, by two independent routes. #1825 fixed the collection side; this is about what happens to the events afterwards.

Up front, the limitation in this report: I cannot demonstrate the runtime symptom. Our fleet is GB200 NVL72, where dcgmi discovery -l reports 0 NvSwitches found on every compute node, so the NVSwitch branch never executes for us and we have zero NVSwitch events to observe. Everything below is from reading the code on main, and the inference about runtime behaviour is labelled as such. Someone on an HGX or DGX baseboard with on-board NVSwitch can confirm it in a minute; I cannot.

1. STORE_ONLY is hardcoded at all three NVSwitch emit sites

health-monitors/gpu-health-monitor/gpu_health_monitor/platform_connector/platform_connector.py, lines 632, 658 and 683, each paired with componentClass="NVSWITCH" at 622, 647 and 676:

processingStrategy=platformconnector_pb2.STORE_ONLY,

commented "NVSwitch remediation is deferred until downstream handling is safe".

STORE_ONLY events are filtered before the Prometheus connector counts them, so they never appear in health_events_total. That much is a deliberate choice about remediation, and I am not arguing with deferring remediation. The side effect is that it also removes them from the counter that alerting is built on, which is a different concern from whether anything should act on them.

Comparable checks expose this as a value. gpuPowerBrakeStoreOnly lets an operator choose observe-only versus acted-on for the power brake. The NVSwitch path has no equivalent knob.

2. The NVSwitch branch never populates dcgm_health_active_events

This is the one I would fix first, because it looks like an oversight rather than a decision, and because it changes no remediation semantics at all.

That gauge is written from pending_metric_updates, drained at line 699:

for event_type, gpu_id, error_code, value in pending_metric_updates:
    metrics.dcgm_health_active_events.labels(...).set(value)

Every append to that list is in the GPU branch, at lines 501, 535 and 588. The NVSwitch branch (roughly 600 to 690) appends to health_events and to pending_cache_updates, and never to pending_metric_updates.

So the one metric that survives STORE_ONLY has no NVSwitch series either.

Combined effect

Inference, not observation: on a system where DCGM does enumerate switches, a DCGM_FR_NVSWITCH_FATAL_ERROR is detected, cached, and written to MongoDB, and is then invisible to health_events_total (filtered as STORE_ONLY) and to dcgm_health_active_events (never recorded). No alert is possible from either metric, and an operator watching a dashboard sees nothing.

That the fault is correctly attributed after #1825 makes this more noticeable, not less: the data is now right and still cannot be seen.

Suggested fixes, in the order I would do them

  1. Append to pending_metric_updates in the NVSwitch branch, mirroring the GPU branch. Two notes from having made the same class of change in feat: label dcgm_health_active_events with the error code #1874: the clear path has to zero the same label set that was set, or a series latches at 1 for the process lifetime; and since feat: label dcgm_health_active_events with the error code #1874 the gauge carries error_code, so a check latched by several codes needs each one cleared.
  2. Make the NVSwitch STORE_ONLY configurable, in the shape of gpuPowerBrakeStoreOnly. Detection-only deployments want these events counted while remediation stays off. I have lower confidence this is wanted, given the comment reads as deliberate, which is why it is second and separable.

Happy to write (1) if it is welcome, with the caveat that I cannot test it on our hardware and would be relying on CI plus a reviewer with switches.

Not duplicates, for the record

Environment

NVSentinel v1.24.0, chart v1.24.0, MongoDB datastore, detection-only. Line numbers verified against origin/main at the time of filing and are identical to the v1.24.0 tag.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions