Is your feature request related to a problem? Please describe.
Some cluster or workload configuration issues can be detected before starting a full-scale training run. These issues may not cause explicit errors, but can silently degrade performance and lead to wasted compute. Examples include incorrect CPU pinning, straggler nodes, lower-than-expected NCCL bandwidth, low filesystem read/write throughput (which can affect data loading, container startup, and checkpointing), incorrect InfiniBand (IB) configuration, an incompatible C++ compiler or toolchain, and tensor-parallel groups that unintentionally span multiple nodes—for example, when TP is greater than the number of GPUs per node.
Tag @NVIDIA/mcore-oncall
to get oncall's attention to this issue.
Describe the solution you'd like
A set of optional preflight checks that run before the training job starts, detect common configuration or infrastructure issues, and report any problems. The checks should provide actionable warnings and, optionally, prevent the job from proceeding when a critical issue is detected.
Describe alternatives you've considered
A default alternative to automated preflight checks is to run experiments designed to measure throughput and compare the results with established performance recipes. Another option is to estimate the upper bound of training throughput, compare it with the observed throughput, and adjust the configuration as needed.
These approaches require additional manual effort and may consume substantial compute before an issue is identified.
Additional context
The checks could cover both static configuration validation and short runtime benchmarks. Where appropriate, users should be able to configure thresholds and choose whether a failed check produces a warning or stops the job.
Is your feature request related to a problem? Please describe.
Some cluster or workload configuration issues can be detected before starting a full-scale training run. These issues may not cause explicit errors, but can silently degrade performance and lead to wasted compute. Examples include incorrect CPU pinning, straggler nodes, lower-than-expected NCCL bandwidth, low filesystem read/write throughput (which can affect data loading, container startup, and checkpointing), incorrect InfiniBand (IB) configuration, an incompatible C++ compiler or toolchain, and tensor-parallel groups that unintentionally span multiple nodes—for example, when TP is greater than the number of GPUs per node.
Tag @NVIDIA/mcore-oncall
to get oncall's attention to this issue.
Describe the solution you'd like
A set of optional preflight checks that run before the training job starts, detect common configuration or infrastructure issues, and report any problems. The checks should provide actionable warnings and, optionally, prevent the job from proceeding when a critical issue is detected.
Describe alternatives you've considered
A default alternative to automated preflight checks is to run experiments designed to measure throughput and compare the results with established performance recipes. Another option is to estimate the upper bound of training throughput, compare it with the observed throughput, and adjust the configuration as needed.
These approaches require additional manual effort and may consume substantial compute before an issue is identified.
Additional context
The checks could cover both static configuration validation and short runtime benchmarks. Where appropriate, users should be able to configure thresholds and choose whether a failed check produces a warning or stops the job.