Skip to content

Measure MFSDP v2 BlockAtomic padding overhead and evaluate row-atomic sharding #7585

Description

@wujingyue

Follow-up to #7114.

MXFP8 groups currently use BlockAtomic(32) to keep shard boundaries aligned to quantization blocks. The shared weight and gradient layout includes padding to satisfy this alignment and equal-sized shards. We should quantify its cost on representative model shapes and data-parallel sizes.

  • Report padding bytes and percentages per parameter group and for the full model, including main weights, gradients, and quantized data and scale planes.
  • Distinguish block-alignment padding from Transformer Engine's physical scale padding.
  • Measure the resulting memory and communication overhead, including configurations with uneven or empty parameter shards.

Based on the measurements, evaluate whether row-atomic sharding could reduce overhead enough to justify its complexity. Evaluate adapting v1's row-atomic approach, which uses an extra all-reduce for the block rowmax. Compare the padding savings against the added communication cost.

Compare with retaining block-atomic boundaries and using uneven communication, tracked in #6587. The outcome should be measurements and a recommendation; a sharding change depends on the results.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions