Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
7f3621e
feat(p2p): add direct RDMA reads into GPU KV cache
GentleCold Sep 11, 2026
4db558c
test(server): preserve default query mode in test requests
GentleCold Sep 11, 2026
c8ac11b
refactor(transfer): propagate GPU registration insertion errors
GentleCold Sep 11, 2026
ec4f4ae
chore(p2p): satisfy direct GPU Clippy checks
GentleCold Sep 11, 2026
8385438
fix(p2p): repair IPC GPU registration and KV block mapping
GentleCold Sep 16, 2026
72b29f7
fix(p2p): preserve RAM prefix before direct GPU fetch
GentleCold Sep 17, 2026
7c00644
refactor(p2p): trim direct GPU RDMA feature scope
GentleCold Sep 17, 2026
20b1c97
fix(p2p): harden direct GPU RDMA failure cleanup
GentleCold Sep 21, 2026
28f7473
fix(p2p): remove unused worker import
GentleCold Sep 21, 2026
f8fd427
fix(p2p): reset RC QP after completion errors
GentleCold Sep 21, 2026
db0d337
fix(p2p): satisfy clippy after QP reset
GentleCold Sep 21, 2026
f513c3a
feat(connector): configure direct GPU RDMA via extra config
GentleCold Sep 22, 2026
9f648ce
perf(direct-gpu): reduce transfer dispatch overhead
GentleCold Sep 23, 2026
af5f0e5
perf(transfer): combine local memory lookups
GentleCold Sep 23, 2026
267f084
perf(rdma): coalesce adjacent read operations
GentleCold Sep 23, 2026
a5c1be6
Revert "perf(rdma): coalesce adjacent read operations"
GentleCold Sep 23, 2026
7559fdb
perf(rdma): batch more read work requests
GentleCold Sep 23, 2026
d0d5f7d
perf(rdma): preallocate completion tracking
GentleCold Sep 23, 2026
188b47c
perf(rdma): add direct GPU fetch stage timing
GentleCold Sep 24, 2026
9ab18ce
perf(rdma): warm direct GPU peer connections
GentleCold Sep 24, 2026
08d856e
Revert "perf(rdma): add direct GPU fetch stage timing"
GentleCold Sep 24, 2026
22da5df
perf(rdma): signal read chains at completion boundaries
GentleCold Sep 24, 2026
31fa000
chore: sync direct GPU RDMA branch with master
GentleCold Sep 24, 2026
4bc9d7e
refactor(p2p): simplify direct GPU RDMA path
GentleCold Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 12 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,9 +54,20 @@ pegaflow-server

```bash
vllm serve Qwen/Qwen3-0.6B \
--kv-transfer-config '{"kv_connector": "PegaKVConnector", "kv_role": "kv_both", "kv_connector_module_path": "pegaflow.connector"}'
--kv-transfer-config '{
"kv_connector": "PegaKVConnector",
"kv_role": "kv_both",
"kv_connector_module_path": "pegaflow.connector"
}'
```

To enable direct GPU P2P reads, add
`"kv_connector_extra_config": {"pegaflow.direct_gpu_rdma": true}` to that
JSON. When it is omitted or set to `false`, remote reads use the host-staging
path. Direct mode currently supports dense attention cache group 0 only and
does not fall back to staging if a direct transfer fails. The same setting can
be passed to `examples/run_vllm_with_pega.py` with `--direct-gpu-rdma`.

> For full server options, multi-node setup, and advanced configuration, see [Server Configuration](./docs/server.md).

## Development
Expand Down
7 changes: 7 additions & 0 deletions docs/deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,13 @@ and use a P/D-aware NIXL router. Configure each server's routable `--addr`,
P and D must use compatible model, tokenizer, block size, KV dtype, KV layout,
and `PYTHONHASHSEED`.

The PegaKVConnector option `pegaflow.direct_gpu_rdma` is configured inside the
connector's `kv_connector_extra_config`. Set it to `true` only on a
`read_write` PegaFlow connector that will load remote blocks; a `save_only`
connector has no PegaFlow load to accelerate. The option is omitted below, so
this deployment uses the host-staging path. See [P2P](./p2p.md#direct-gpu-performance-notes)
for the direct-load command and the benchmark caveats.

Replace `<p_node_ip>` and `<d_node_ip>` with the addresses assigned to the P
and D nodes. The example assumes P and D run on separate nodes; to colocate
them on one host, keep the distinct NIXL side-channel ports and point both
Expand Down
33 changes: 32 additions & 1 deletion docs/metrics.md
Original file line number Diff line number Diff line change
Expand Up @@ -193,7 +193,7 @@ The setting remains configurable with `--metric-hll-bucket-bits`.
### Load Metrics (CPU → GPU)
- **pegaflow_load_bytes_total** (Counter)
- Total bytes loaded from CPU storage to GPU
- Use case: Monitor load throughput
- Use case: Monitor host-staging load throughput

- **pegaflow_load_duration_seconds** (Histogram)
- Load operation latency distribution
Expand All @@ -203,6 +203,36 @@ The setting remains configurable with `--metric-hll-bucket-bits`.
- Load operation failures (e.g., transfer errors)
- Use case: Detect data transfer issues

### Direct GPU RDMA Load Metrics

These metrics cover the opt-in `pegaflow.direct_gpu_rdma` path. A direct
remote read bypasses requester host staging, so it is not included in
`pegaflow_load_bytes_total` or `pegaflow_load_duration_seconds`. A mixed load
can still emit the ordinary load metrics for its local RAM prefix while the
remote suffix is represented by the direct metrics below.

- **pegaflow_direct_gpu_load_total** (Counter, `status=success|error`)
- Number of direct GPU load batches reaching a terminal state
- Use case: Confirm that the direct path was exercised and detect failures

- **pegaflow_direct_gpu_load_duration_seconds** (Histogram)
- End-to-end direct load batch duration
- Includes remote block metadata queries, RDMA completion, GPU visibility
flush, and any concurrent local H2D work
- Use case: Compare direct and host-staging load latency for the same hit set

- **pegaflow_direct_gpu_mr_registration_failures** (Counter)
- CUDA DMA-BUF or RDMA memory-registration failures
- Use case: Detect capability or allocation-registration problems before a
direct load can start

The current direct path does not export direct-RDMA bytes or per-stage
histograms. Do not compare `pegaflow_direct_gpu_load_duration_seconds` with
`pegaflow_rdma_fetch_duration` as if they had identical boundaries, and do not
use `pegaflow_rdma_fetch_bytes` to infer direct-GPU bandwidth. Use the direct
load counters together with request-level timings and debug logs for an A/B
comparison.

### SSD Cache Metrics
- **pegaflow_ssd_write_bytes_total** (Counter) - Bytes written to SSD cache
- **pegaflow_ssd_write_duration_seconds** (Histogram) - SSD write latency
Expand All @@ -229,6 +259,7 @@ This metric intentionally records decisions, not completed service outcomes.
For backing failure correlation, use:

- `pegaflow_rdma_fetch_total{status="error"}` for RDMA fetch failures
- `pegaflow_direct_gpu_load_total{status="error"}` for direct GPU load failures
- `pegaflow_ssd_prefetch_failures_total` for SSD prefetch failures

The legacy `pegaflow_cache_block_hits_total` and
Expand Down
93 changes: 90 additions & 3 deletions docs/p2p.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,27 +81,108 @@ pegaflow-server \

### 3. Launch inference engine

Same as single-node — PegaFlow server handles P2P transparently.
Same as single-node — PegaFlow server handles host-staged P2P transparently.

```bash
vllm serve Qwen/Qwen3-0.6B \
--kv-transfer-config '{"kv_connector": "PegaKVConnector", "kv_role": "kv_both", "kv_connector_module_path": "pegaflow.connector"}'
```

To select direct GPU reads, pass the connector option through vLLM's
`kv_connector_extra_config`:

```bash
vllm serve Qwen/Qwen3-0.6B \
--kv-transfer-config '{
"kv_connector": "PegaKVConnector",
"kv_role": "kv_both",
"kv_connector_module_path": "pegaflow.connector",
"kv_connector_extra_config": {
"pegaflow.direct_gpu_rdma": true
}
}'
```

When enabled, the scheduler carries the remote fetch plan in the query lease
and the worker issues RDMA READs into vLLM's registered GPU KV allocations.
Omitting the option (or setting it to `false`) keeps the existing path of RDMA
READ into host memory followed by host-to-GPU copy. Direct mode currently
supports dense attention cache group 0 only. A failed direct load is surfaced
to vLLM so the request can recompute; it is not automatically retried through
host staging.

### 4. Verify

Use `--log-level debug` to confirm P2P is working. Look for MetaServer registration, RDMA handshake, and RDMA fetch messages in the logs.

## Fallback Behavior
## Failure behavior

P2P is opportunistic. Failures degrade gracefully to single-node operation — no crashes, no significant performance impact in most cases.
Host-staged P2P is opportunistic and can proceed without a remote hit when
discovery or transfer fails. Direct GPU mode has a stricter failure boundary:
the failed load is reported to vLLM and the affected prefix is recomputed.

| Scenario | What happens |
|---|---|
| MetaServer unreachable | Hash registration silently dropped. No remote discovery attempted. |
| Remote node unreachable | gRPC handshake fails, fetch aborted. Request proceeds without remote blocks. |
| RDMA transfer timeout | Connection invalidated, transfer lock force-released. Logged as error. |

For direct GPU mode, the last two cases fail the direct load and do not invoke
the host-staging path.

## Direct GPU performance notes

Direct GPU mode removes the requester-side host-to-GPU copy. It does not remove
the other work on the critical path: MetaServer query, the per-segment gRPC
block-metadata query, connection setup on a cold peer, descriptor construction,
RDMA completion waits, transfer-lock release, and the CUDA GPUDirect visibility
flush. The current implementation processes segments in order and flushes the
CUDA context after each segment, so short prefixes or many owner segments can
be latency-bound even when the raw GB300 GPU RDMA bandwidth is higher than the
host-staging path.

The existing 2P2D replay is not an apples-to-apples direct-load benchmark when
NIXL is configured as the P-to-D transfer connector. NIXL performs the main
P-to-D GPU transfer, while PegaFlow only handles the P-side cache lookup/load
for the requests that hit a remote PegaFlow owner. The decode-side
`save_only` connector records no PegaFlow loads. A small end-to-end difference
in that replay therefore does not measure the direct-vs-host-staging copy in
isolation.

The GB300 replay illustrates the dilution. In one paired 8,422-request run,
direct mode completed 1,387 direct loads in 16.322 seconds in aggregate
(about 11.8 ms per direct load), while the host-staging control completed 735
host RDMA fetches in 9.607 seconds (about 13.1 ms per fetch) and had one fetch
error. The direct and host counters cover different cache histories and
different operation boundaries, so these averages are directional evidence,
not a throughput ratio. Overall mean TTFT was 187.527 ms versus 187.460 ms and
throughput was 5.322 versus 5.326 requests/s. The run therefore shows why a
faster direct data path can produce little end-to-end movement: most requests
are misses, and the remaining hit latency includes scheduler work, P-side
prefill, NIXL P-to-D transfer, and decode startup.

To compare the two PegaFlow paths, use the same model, block count, request
ordering, remote owner, and hit set. Make PegaKVConnector the requester-side
`read_write` connector and run the same requests once with direct mode and
once with host staging. For direct mode, use
`pegaflow_direct_gpu_load_total` and
`pegaflow_direct_gpu_load_duration_seconds`; for host staging, use the
`pegaflow_rdma_fetch_*` metrics. The latter do not include direct GPU loads,
and the current direct path does not export a direct-RDMA byte counter, so do
not combine those counters into one bandwidth number. Split the measurement
with connector timings and debug logs into query, metadata/handshake, RDMA
transfer, visibility flush, and vLLM scheduling time.

The current direct implementation has four likely latency costs after the
query: one metadata RPC per owner segment, sequential segment processing,
descriptor construction and completion waits, and creating/binding a CUDA
context plus a GPUDirect visibility flush after each segment. Connection reuse
removes the handshake from warm peers, and local RAM H2D and remote GPU RDMA
already run concurrently. The first optimization candidates are therefore
coalescing or parallelizing owner segments and moving the visibility flush to
the end of a load when the CUDA/RDMA contract permits it; measure each change
with the isolated A/B workload above before changing the data path.

## Tuning

### Hugepages
Expand Down Expand Up @@ -137,6 +218,9 @@ P2P-related Prometheus metrics (on `:9091/metrics` by default):
| `pegaflow_rdma_fetch_bytes` | Counter | Total bytes fetched via RDMA |
| `pegaflow_rdma_fetch_plan_segments` | Histogram | Planned segment count per executed RDMA fetch plan |
| `pegaflow_rdma_fetch_plan_completed_segments` | Histogram | Completed segment count before a plan stops |
| `pegaflow_direct_gpu_load_total` | Counter | Direct GPU load attempts, labelled by success or error |
| `pegaflow_direct_gpu_load_duration_seconds` | Histogram | End-to-end direct GPU load duration |
| `pegaflow_direct_gpu_mr_registration_failures` | Counter | GPU memory registration failures that prevent direct loads |
| `pegaflow_rdma_qps` | Gauge | Active RDMA queue pairs |
| `pegaflow_transfer_lock_active` | UpDownCounter | Currently held transfer locks |
| `pegaflow_transfer_lock_timeouts_total` | Counter | Transfer lock timeout events |
Expand All @@ -155,3 +239,6 @@ P2P-related Prometheus metrics (on `:9091/metrics` by default):
- Enable hugepages for large pools (`--use-hugepages`).

For all P2P issues, `--log-level debug` shows the full handshake and fetch flow.
The direct load duration includes the remote query, RDMA completion, GPU
visibility flush, and any concurrent local H2D load. It is not a raw GPU-RDMA
bandwidth measurement.
13 changes: 13 additions & 0 deletions examples/run_vllm_with_pega.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,14 @@ def parse_args():
default=0.9,
help="GPU memory utilization (default: 0.9)",
)
parser.add_argument(
"--direct-gpu-rdma",
action="store_true",
help=(
"Read remote KV blocks directly into GPU memory with GPUDirect RDMA "
"(default: disabled)"
),
)
parser.add_argument(
"--kv-events",
action="store_true",
Expand Down Expand Up @@ -77,6 +85,10 @@ def main():
"kv_role": "kv_both", # Both scheduler and worker roles
"kv_connector_module_path": "pegaflow.connector",
}
if args.direct_gpu_rdma:
kv_transfer_config["kv_connector_extra_config"] = {
"pegaflow.direct_gpu_rdma": True
}

# Build vllm serve command
cmd = [
Expand Down Expand Up @@ -114,6 +126,7 @@ def main():
print(f"Model: {args.model}")
print(f"Endpoint: http://{args.host}:{args.port}")
print(f"Tensor Parallel Size: {args.tensor_parallel_size}")
print(f"Direct GPU RDMA: {'enabled' if args.direct_gpu_rdma else 'disabled'}")
if args.kv_events:
print("Prefix Caching: enabled (required for KV events)")
print(
Expand Down
2 changes: 1 addition & 1 deletion pegaflow-core/src/backing/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ use pegaflow_common::NumaNode;
#[cfg(feature = "rdma")]
pub(crate) use rdma::{RdmaTransport, new_rdma};
#[cfg(feature = "rdma")]
pub(crate) use rdma_fetch::RdmaFetchStore;
pub(crate) use rdma_fetch::{DirectFetchPlan, DirectQueryPlan, GpuReadTarget, RdmaFetchStore};
pub(crate) use ssd::SsdBackingStore;
pub(crate) use ssd::new_ssd;

Expand Down
Loading
Loading