Scaling Bitwise-Deterministic Pretraining for a Trillion-Parameter Nemotron Model with NVIDIA Megatron Core #7497
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Authors: @euronymous-aithal, @ericharper, @sbhavani, @Connor-XY, @ZhiyuLi-Nvidia
Debugging large language model training becomes difficult and expensive at trillion-parameter scale across thousands of GPUs, especially when multiple forms of parallelism, low-precision computation, and distributed checkpointing interact. Reproducing issues is more difficult if the training is non-deterministic since the issue may or may not occur in repeated runs. Validating fixes to the issues also remains challenging since the improved numerics could come from variance rather than the bug fix when training with non-determinism.
Production and hero training runs are expensive and require a strong guarantee that they will be successful. Bitwise determinism is essential for reliably replaying failures, debugging loss spikes, validating system changes, and recovering interrupted training runs. Due to the cost, large-scale training runs must also be extremely efficient. At “megatron scale”, even a small slowdown can cost thousands of GPU-days, making performance optimization essential for enabling determinism while training.
Using a Trillion-Parameter Nemotron Model pretraining as a case study, this post explains how NVIDIA is developing bitwise determinism as an end-to-end capability in Megatron Core. We will cover:
What we mean by bitwise determinism
Given fixed inputs, independent runs must follow the same numerical trajectory, and saving and resuming from a checkpoint must not change that trajectory.
“Fixed inputs” encompass more than a random seed. They include:
Megatron Core determinism targets two guarantees:
Independent reproducibility
Two runs launched from the same initial state must remain bitwise identical from beginning to end.
Here, “bitwise identical” means that two runs follow exactly the same training trajectory, step by step and bit for bit.
Bitwise checkpoint resume
A run that saves and restores one or more checkpoints must remain identical to an uninterrupted run.
The objective is one training trajectory.
Why pretraining should be deterministic
Replaying critical training events
If a loss spike occurs during a nondeterministic run, restarting the job may produce a different trajectory and make the event disappear. Engineers are then left investigating an incident that they cannot reproduce.
With deterministic replay, the spike reappears at the same step. Components can then be changed one at a time until the spike moves or disappears, providing a controlled way to isolate the cause.
Preventing checkpoint-induced trajectory changes
Large training jobs are routinely interrupted by planned maintenance, infrastructure failures, or scheduling requirements. Without bitwise checkpoint resume, restoring the job can alter its random-number stream, data position, low-precision weights, or optimizer state.
The resumed job may remain numerically healthy, but it is no longer the same experiment.
Detecting silent corruption
Once the software stack is deterministic, a mismatch between two repeated runs becomes a strong diagnostic signal.
This can expose silent checkpoint corruption, unexpected software behavior, or potentially faulty hardware. A deterministic recipe is then both a regression test and a workload for testing fleet health.
Making A/B comparisons more trustworthy
A deterministic reference improves controlled experiments. When two runs differ by exactly one intended change, any numerical divergence can be attributed to that change rather than background execution noise.
How to verify determinism
Printed loss values are not sensitive enough to establish bitwise identical numerics. Two runs can print the same rounded loss while already differing in the lower-order bits of their gradients or parameters.
A stronger validation procedure fingerprints the numerical state at every step. Depending on the investigation, this can include:
We run the same configuration twice and compare these fingerprints step by step. The first mismatching step helps to indicate the root cause of the underlying issue.
In order to validate bitwise checkpoint resume we compare a continuous reference run with one or more interrupted runs. After every restore, verify that the parameters, optimizer state, RNG state, data position, and subsequent outputs are bitwise identical to the reference.
How to fix determinism when it breaks
Nondeterminism in large-scale training workloads may be intermittent, appear only at high GPU counts, or depend on a particular parallelism configuration. A recipe that is bitwise exact at small scale can diverge at production scale.
The point where two loss curves separate rarely identifies where nondeterminism began. It only shows where the difference became large enough to appear in the logged metric. The first differing bit may have occurred much earlier.
The agent skill proposed in Megatron-LM PR #7262 records ordered, per-rank streams of tensor fingerprints and compares two runs offline. The PR documents a diagnostic workflow and implementation guidance.
Use a repeatable workflow: locate the first divergent record, determine whether its inputs matched, identify the underlying mechanism, and validate the correction with paired runs.
Before tracing, confirm that both runs use the same seed, data order, global batch size, parallelism layout, container, and software stack. Compare full-precision serialized metrics rather than rounded console output.
Progressively localize the divergence with granular tracing
A determinism break can originate anywhere in the training stack. Begin with an end-to-end comparison, then progressively narrow the capture scope from the training iteration to the phase, module, operation, and kernel.
At every level, apply the same test: if two runs enter a scope with bitwise-identical inputs but leave it with different outputs, the divergence originated within that scope.
The investigation begins with broad, low-overhead coverage and becomes more granular only as the search space contracts. Do not trace only the iteration where the loss visibly separates; the first differing bit may have appeared earlier. Use broad tracing to locate the earliest divergent iteration and ranks, then enable detailed operation and kernel tracing only within that narrowed scope.
Compare traces offline
Each selected rank writes an append-only file stream without adding collectives or cross-rank ordering to the observed training step. This avoids changing execution timing or masking the race being investigated.
We align events by a run-independent identity, such as the operation name, occurrence count, and module scope.
For each aligned record, the main test is:
Matching inputs with different outputs identify a candidate origin. If both inputs and outputs differ, the operation received an upstream difference, so continue walking backward.
The word “first” is exact only within one rank. Ranks do not share a global operation clock, so their sequence numbers cannot be compared directly. Classify each rank’s first mismatch causally:
If many ranks identify the same originating operation, the operation itself is likely nondeterministic. If only a subset does, we need to further investigate likely causes such as topology, rank placement, input distribution, or reduction ordering.
Fingerprint tensors on the GPU
The proposed workflow recommends
torch.hash_tensorfor fast, GPU-resident fingerprints. Record each tensor’s shape, dtype, and element count alongside its digest. For MXFP8 or NVFP4 tensors, fingerprint both the encoded values and their scale buffers.Whole-tensor XOR fingerprints cannot detect permutations, which is important for routing maps and MoE dispatch outputs. Fingerprint these tensors by row or chunk using the
dimargument.A fingerprint is an efficient check, but does not guarantee that the tensors are bitwise identical. Use stronger byte-level comparisons for confirming collisions.
Rule out false alarms
Three cases can make a correct trace point to the wrong cause:
TorchDispatchModeobserves ATen operations routed through the PyTorch dispatcher. Custom kernels can bypassTorchDispatchModeand escape the tracing scope. If the first mismatch appears at a simple view, slice, or addition, probe the custom kernel that produced its input.Validate the fix
Validate a patch with paired runs:
Note, bitwise determinism is validated within the same hardware and software environment. Comparisons across GPU generations, network configurations, or library versions are outside this guarantee and may produce different numerical results.
Optimizing a deterministic training of a Trillion-Parameter Nemotron Model
Correctness alone is not enough for production determinism. A deterministic recipe with substantial overhead may support debugging, but it is not efficient enough to adopt for a hero run. Deterministic and nondeterministic execution should be co-optimized from the beginning.
Model Architecture of Nemotron Model
This work uses a trillion-parameter Nemotron model as a representative production-scale workload. The model combines Mamba-style state-space model (SSM) layers with Transformer attention layers in a hybrid architecture.
The SSM path performs sequence mixing through recurrent state updates, scans, and convolutional operations. The Transformer path provides attention-based token interactions. The hybrid design combines these complementary layer types within one training recipe.
This architecture also broadens the determinism surface. Bitwise reproducibility must hold across SSM kernels, attention kernels, low-precision computation, distributed communication, and the transitions between different layer types.
Establish a controlled baseline
Compare deterministic execution with the fastest supported nondeterministic recipe using the same model, hardware, batch sizes, parallelism strategy, precision format, software environment, and measurement window.
Record:
Calculate the overhead as:
Note we should start collecting the data points when the training performance is stable, as data points taken before will skew the results.
Optimization journey
Performance optimization began with Nemotron 3 Ultra. The initial MCore baseline had an approximately 15% determinism tax on 96 GPUs. Avoiding the fill of uninitialized memory reduced the tax to approximately 1–2%. At 1,536 GPUs, the large-scale proxy measured an approximately 1.5% determinism tax.
The early Nemotron Triton recipe had an approximately 38% determinism tax on eight GPUs. Restoring MoE-MLP fusion reduced the tax to approximately 21%. Caching the autotune configuration and capping
num_warpsfurther reduced it to approximately 2%.The Nemotron CuteDSL baseline initially broke determinism at 256 GPUs. After work on the MoE-MLP weight-gradient path, the determinism tax was within measurement noise at 256 and 512 GPUs and approximately 3.5% at 1,024 GPUs. The large-scale Nemotron recipe measured an approximately 4.6% steady-state determinism tax at 3,072 GPUs.
These results show that determinism performance must be addressed at several levels. Memory handling, fusion, autotuning, kernel configuration, and the weight-gradient path each affected the final overhead.
Determinism performance optimization milestones
num_warpsExample of Kernel Level Optimization
The following are three general concepts for writing an optimized deterministic kernel:
The grouped-GEMM epilogue provides a concrete example of this approach. Multiple N-tiles originally accumulated into the same
dprob[token]address, making the result depend on their arrival order. Serializing those writers restored determinism but reduced parallelism.The optimized solution is to separate writers in space, then order once. Give each N-tile a private output slot, preserve parallel execution inside the kernel, and combine the slots in a fixed order after all writers finish.
Validate optimization at scale
A low determinism tax on a few GPUs does not guarantee the same result at production scale. Communication, synchronization, pipeline bubbles, expert routing, and load balance can change the relative overhead.
The optimized recipe must therefore be validated at progressively larger GPU counts. Small-scale measurements provide a fast optimization loop, while large-scale measurements determine whether the result is suitable for the hero run.
Impact of the determinism tax
For a hypothetical 100-day run using 10,000 GPUs:
Maintain determinism as models, kernels, and recipes evolve
Over time non-determinism can be introduced through: new kernels, fusions, precision formats, or parallelism configurations. Megatron-LM now protects against possible determinism breaks at several different levels:
--deterministic-modeapplies canonical environment settings, enables PyTorch deterministic algorithms, and rejects features without a deterministic path.Together, these checks help maintain Megatron-LM as a deterministic framework. The current coverage and known gaps are kept up to date in: determinism status, operation catalog, and kernel testing guide.
Common sources of nondeterminism
Runtime autotuning selects different kernels
Autotuners often benchmark multiple kernel configurations at startup and retain the fastest result. Timing noise can cause different ranks or repeated runs to select different tile shapes.
Different shapes may change the order of floating-point accumulation, producing different bits from the first training step.
For deterministic execution, the selected configuration must be fixed or reliably cached.
A fused reduction uses an unstable accumulation order
Floating-point addition is not associative. Changing the order of a reduction can change its final bits.
A fused loss or gradient kernel that uses atomics or an unordered reduction may produce different results even when its inputs match exactly. One safe interim solution is to disable that fusion in deterministic mode while developing a deterministic implementation with a fixed reduction strategy.
Checkpoint resume restores an incomplete RNG state
If a checkpoint omits any RNG state, the resumed run begins consuming a different random stream. That can affect dropout masks, data shuffling, MoE routing behavior, or initialization performed after the restore.
A deterministic checkpoint must save and restore every relevant random-number state bit for bit.
A low-level library kernel changes the result
The framework may be configured correctly while a single underlying library kernel remains nondeterministic.
Tracing can isolate the affected operation so deterministic mode can steer around it while a focused reproducer is supplied to the owning library team.
Checkpoints must preserve low-precision state
Checkpoint correctness becomes more subtle when the live training representation uses low-precision weights plus scaling metadata.
Suppose a quantized weight is saved in a different precision while its original scale is discarded. Loading the checkpoint then requires requantization and scale reconstruction. Even if the difference is small enough to leave the visible loss curve unchanged, the resumed model is not bitwise identical.
The trillion-parameter Nemotron model checkpoint path addresses this by preserving a full-precision source of truth and reconstructing the runtime representation through the same conversion path used during training.
The governing rule is strict: saving and loading a checkpoint must not change a single tracked bit.
Known operations without determinism support
Some operations do not yet have a deterministic implementation. These are known support limitations rather than newly introduced determinism regressions. A training recipe that uses one of these operations may require a deterministic alternative or a configuration that avoids the unsupported path.
Consult the Megatron Core catalog of operations without determinism support for current operator-level coverage gaps and constraints.
All reactions