Skip to content

TransferBench v1.71.00 - #359

Open
AtlantaPepsi wants to merge 13 commits into
developfrom
candidate-1.71
Open

AtlantaPepsi wants to merge 13 commits into
developfrom
candidate-1.71

Conversation

@AtlantaPepsi

@AtlantaPepsi AtlantaPepsi commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

  • NIC executor check against max message size
  • fixing logical CU ID report
  • inclusion of gfx1250-strict target
  • Pingpong latency testing integration and presets

Technical Details

Test Plan

  • Validate HSA_DISABLE_GFX12_STRICT=0 ./TransferBench output of gfx1250-strict target
  • On gfx1250, gfx942 and gfx950 compare SHOW_ITERATIONS output CU ID against CU_MASK bits
  • Latency Testing
    • p2p_latency preset: in/cross-domain, and AMD and Nvidia platform
    • Concurrent transfers + pingpongs run, single/multistream
    • validation of all memory types support

Test Result

p2p_latency example output

[Latency Related]
GPU_MEM_TYPE         =            0 : Using default GPU memory for flags (0=default, 1=fine-grained, 2=uncached, 3=managed)
NUM_GPU_DEVICES      =            8 : Using 8 GPUs
NUM_LAPS             =         1000 : Timing 1000 round trips per iteration
USE_REMOTE_READ      =            0 : Executors write to their partner's memory and poll their own

Pingpong round-trip latency per lap (us), each pair run by itself
[1000 laps] [default GPU memory flags] [remote write / local poll]
 PING\PONG     GPU 00     GPU 01     GPU 02     GPU 03     GPU 04     GPU 05     GPU 06     GPU 07
    GPU 00      0.550      1.973      4.093      4.504      4.155      4.501      4.092      4.504
    GPU 01      1.956      0.585      4.502      4.930      4.507      4.860      4.505      4.925
    GPU 02      4.079      4.501      0.548      1.993      4.095      4.503      4.067      4.503
    GPU 03      4.501      4.919      1.995      0.592      4.503      4.873      4.502      5.010
    GPU 04      4.111      4.501      4.050      4.501      0.578      1.973      4.111      4.501
    GPU 05      4.501      4.901      4.501      4.837      1.993      0.581      4.503      4.818
    GPU 06      4.060      4.501      4.034      4.500      4.088      4.501      0.562      1.954
    GPU 07      4.502      4.957      4.502      4.873      4.511      4.841      1.956      0.575

Concurrent transfer + pingpong hybrid run example output

./TransferBench cmdline 16M "6 16 (G0->G1->G1) (G1->G2->G2) (G2->G3->G3) (G3->G0->G0
) (N G0 G1 +1000 N G1 G0) (C0 G0 G1 +50 C0 G1 G0)"

Test 1:
-------------------┬--------------┬------------┬-------------------┬---------------------------
  Executor: GPU 00 │ 302.684 GB/s │   0.055 ms │    16777216 bytes │ 332.617 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 3    │ 332.617 GB/s │   0.050 ms │    16777216 bytes │ G3 -> G0:16 -> G0
     PingPong 4    │     1.046 us │   1.046 ms │         1000 laps │ N->G0->G1 <+> N->G1->G0
     PingPong 5    │     1.544 us │   0.077 ms │           50 laps │ C0->G0->G1 <+> C0->G1->G0
-------------------┼--------------┼------------┼-------------------┼---------------------------
  Executor: GPU 01 │ 331.512 GB/s │   0.051 ms │    16777216 bytes │ 355.359 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 0    │ 355.359 GB/s │   0.047 ms │    16777216 bytes │ G0 -> G1:16 -> G1
-------------------┼--------------┼------------┼-------------------┼---------------------------
  Executor: GPU 02 │ 345.778 GB/s │   0.049 ms │    16777216 bytes │ 350.870 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 1    │ 350.870 GB/s │   0.048 ms │    16777216 bytes │ G1 -> G2:16 -> G2
-------------------┼--------------┼------------┼-------------------┼---------------------------
  Executor: GPU 03 │ 358.855 GB/s │   0.047 ms │    16777216 bytes │ 364.532 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 2    │ 364.532 GB/s │   0.046 ms │    16777216 bytes │ G2 -> G3:16 -> G3
-------------------┼--------------┼------------┼-------------------┼---------------------------
   Aggregate (CPU) │  58.392 GB/s │   1.149 ms │    67108864 bytes │ Overhead 1.094 ms
-------------------┴--------------┴------------┴-------------------┴---------------------------

Submission Checklist

Copilot AI lite review requested due to automatic review settings September 23, 2026 15:27
@AtlantaPepsi
AtlantaPepsi requested review from a team as code owners September 23, 2026 15:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved critical and moderate correctness, portability, parsing, and build-target issues remain.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 4 High severity · 4 Medium severity

Open (8)
What changed in this PR

Adds pingpong latency testing, NIC message-size validation, logical CU reporting fixes, and gfx1250-strict support.

Changes:

  • Adds pingpong parsing, execution, timing, and latency presets.
  • Adds NIC max_msg_sz validation and reporting.
  • Updates GPU architecture and TDM handling.
File Review summary
src/​header/​TransferBench.hpp Critical and moderate issues in XCC handling, pingpong validation, NIC limits, subindex selection, dump parsing, and timing scaling.
src/​header/​tdmCopy.h Critical architecture guard mismatch enables unsupported TDM targets.
src/​client/​Utilities.hpp Reviewed result and topology utilities.
src/​client/​Topology.hpp Reviewed NIC message-size reporting.
src/​client/​Presets/​Presets.hpp Reviewed latency preset registration.
src/​client/​Presets/​Latency.hpp Moderate issues with failed-run result handling and diagonal pair measurement.
src/​client/​EnvVars.hpp Reviewed pingpong configuration variables.
src/​client/​Client.cpp Reviewed pingpong transfer display updates.
docs/​install/​build_from_source.rst Reviewed strict GPU target documentation.
CMakeLists.txt Moderate issue: package build targets omit gfx1250-strict.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/header/TransferBench.hpp
exeInfo.totalSubExecs += t.numSubExecs;
} else {
exeInfo.totalPingpong ++;
}
exeInfo.useSubIndices |= (t.exeSubIndex != -1 || (t.exeDevice.exeType == EXE_GPU_GFX && !cfg.gfx.prefXccTable.empty()));
Comment thread src/header/TransferBench.hpp Outdated
Comment thread src/header/tdmCopy.h
Comment thread src/client/Presets/Latency.hpp Outdated
Comment thread src/header/TransferBench.hpp
Comment thread src/header/TransferBench.hpp
Comment on lines +8370 to +8373
auto singleMemOrNull = [](vector<MemDevice> const& mems) {
if (mems.empty()) return MemDevice{MEM_NULL, 0, 0};
return mems[0];
};
Copilot AI review requested due to automatic review settings September 23, 2026 15:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread src/header/TransferBench.hpp
@AtlantaPepsi AtlantaPepsi changed the title Candidate 1.71 TransferBench v1.71.00 Sep 23, 2026
Copilot AI review requested due to automatic review settings September 23, 2026 21:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

#else
useSubIndexCount[exe]++;
int numSubIndices = GetNumExecutorSubIndices(exe);
if (subIndex >= numSubIndices) {
Comment on lines +6660 to +6661
dim3 const gridSize(xccDim, numPingpong, 1);
dim3 const blockSize(1);
Comment on lines 3111 to +3115
if (t.numSubExecs <= 0)
errors.push_back({ERR_FATAL, "Transfer %d: # of subexecutors must be positive", i});
else
else if (isPingpong) {
if (t.numSubExecs != 1)
errors.push_back({ERR_WARN,
Copilot AI review requested due to automatic review settings September 28, 2026 15:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +5562 to +5567
} else if (partnerMem.memIndex != exeDevice.exeIndex) {
if (System::Get().IsVerbose()) {
System::Get().Log("[INFO] Enabling pingpong peer access: GPU %d -> GPU %d\n",
exeDevice.exeIndex, partnerMem.memIndex);
}
ERR_CHECK(EnablePeerAccess(exeDevice.exeIndex, partnerMem.memIndex));
Comment thread src/client/Presets/Latency.hpp
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants