Follow-up to #7114.
MXFP8 groups currently use BlockAtomic(32) to keep shard boundaries aligned to quantization blocks. The shared weight and gradient layout includes padding to satisfy this alignment and equal-sized shards. We should quantify its cost on representative model shapes and data-parallel sizes.
- Report padding bytes and percentages per parameter group and for the full model, including main weights, gradients, and quantized data and scale planes.
- Distinguish block-alignment padding from Transformer Engine's physical scale padding.
- Measure the resulting memory and communication overhead, including configurations with uneven or empty parameter shards.
Based on the measurements, evaluate whether row-atomic sharding could reduce overhead enough to justify its complexity. Evaluate adapting v1's row-atomic approach, which uses an extra all-reduce for the block rowmax. Compare the padding savings against the added communication cost.
Compare with retaining block-atomic boundaries and using uneven communication, tracked in #6587. The outcome should be measurements and a recommendation; a sharding change depends on the results.
Follow-up to #7114.
MXFP8 groups currently use
BlockAtomic(32)to keep shard boundaries aligned to quantization blocks. The shared weight and gradient layout includes padding to satisfy this alignment and equal-sized shards. We should quantify its cost on representative model shapes and data-parallel sizes.Based on the measurements, evaluate whether row-atomic sharding could reduce overhead enough to justify its complexity. Evaluate adapting v1's row-atomic approach, which uses an extra all-reduce for the block rowmax. Compare the padding savings against the added communication cost.
Compare with retaining block-atomic boundaries and using uneven communication, tracked in #6587. The outcome should be measurements and a recommendation; a sharding change depends on the results.