Skip to content

fix(xet): give each read stream its own download buffer (#234) - #241

Draft
XciD wants to merge 2 commits into
mainfrom
fix/per-stream-download-buffer
Draft

XciD wants to merge 2 commits into
mainfrom
fix/per-stream-download-buffer

Conversation

@XciD

@XciD XciD commented Sep 28, 2026

Copy link
Copy Markdown
Member

Fixes #234.

Problem

Under many concurrent cold sequential readers (parallel -j64 sha512sum -c over the mount), reads timed out at cursor=0, retried, and ended in Input/output error. Throughput collapsed to about 1 Gbps on a 200 Gbps link.

Every read stream competed FIFO for xet-core's single global download buffer, which hf-mount capped at 256 MiB. Buffer held by downloaded but unconsumed data is only released when a FUSE worker thread consumes it, and those threads were all blocked waiting on streams that could not get buffer for their first term. The 30 s fetch timeout broke the cycle, but the retry re-entered the FIFO at the back, so it mostly re-downloaded and eventually gave up with EIO. --direct-io only reduced kernel readahead piling onto the same handle; the starvation was the same.

Fix

Each stream gets a private AdjustableSemaphore through FileReconstructor::with_buffer_semaphore, so streams never wait on each other. The size is a fair share of HF_XET_RECONSTRUCTION_DOWNLOAD_BUFFER_LIMIT (now 1 GiB by default) clamped to [one xorb, 256 MiB], rebalanced when a stream opens or closes. A lone reader keeps deep pipelining; many readers each keep a guaranteed floor. In-flight memory is bounded by max(limit, streams * 64 MiB).

The floor cannot go below one xorb (64 MiB): xet-core clamps a term acquire against the semaphore total only when it is issued and does not re-clamp pending acquires on shrink, so a smaller floor left in-flight acquires that could never complete (seen as simultaneous timeouts at cursor=24 MiB during development).

A second commit raises the adaptive download concurrency cap from 64 to 124 (xet-core's high-performance preset value). It is not measured in isolation; the numbers below include it.

Measurements

c7gn.16xlarge (64 vCPU, 200 Gbps, arm64, kernel 6.8), 128 x 1 GiB files, the reporter's command with --max-threads 64 --no-disk-cache, cold cache:

build mode elapsed avg throughput stream timeouts I/O errors
main page cache 996 s 1.2 Gbps 108 0
main --direct-io 1000 s 1.2 Gbps 90 0
fix page cache 59 s 19.3 Gbps (peak 25.7) 0 0

Single sequential reader, cold: 0.8 to 1.1 Gbps on main, 1.2 to 1.35 Gbps with the fix. Daemon peak RSS under the 64-reader run: 7.8 GiB on main, 8.9 GiB with the fix (the per-handle readahead windows dominate in both).

Tests

  • test_fuse_parallel_cold_reads: 24 cold readers on 4 FUSE threads with a 5 s fetch timeout. Fails on main with 13 of 24 readers in EIO and 84 stream timeouts; passes with the fix in about 5 s.
  • Unit tests for the share computation, including a shrink while permits are held.
  • Full fuse_ops suite green on the Linux instance; 420 unit tests green.

Every read stream competed FIFO for xet-core's single global download
buffer, which hf-mount capped at 256 MiB. Buffer held by downloaded but
unconsumed data is only released when a FUSE worker thread consumes it,
and under many concurrent cold readers those threads were all blocked
waiting on streams that could not get their first term. Reads timed out
at cursor=0, retried at the back of the queue, and ended in EIO.

Each stream now gets a private AdjustableSemaphore via
FileReconstructor::with_buffer_semaphore, sized as a fair share of the
download buffer limit (1 GiB by default) clamped to [one xorb, 256 MiB].
The floor cannot go below one xorb: xet-core clamps a term acquire only
when it is issued, so a shrink below an in-flight acquire would leave it
unsatisfiable.

Measured on c7gn.16xlarge, 128 x 1 GiB files, parallel -j64 sha512sum -c:
main 996 s at 1.2 Gbps with 108 stream timeouts; fix 59 s at 19.3 Gbps
with none. Single-stream reads go from 0.8-1.1 Gbps to 1.2-1.35 Gbps.

Adds test_fuse_parallel_cold_reads (24 cold readers on 4 FUSE threads),
which fails on main with 13 of 24 readers in EIO.
With per-stream buffers, many concurrent readers keep more terms in
flight than xet-core's default cap of 64 connections allows. 124 is the
value of xet-core's own high-performance preset. Not measured in
isolation; the #234 bench numbers were taken with this cap.
@github-actions

Copy link
Copy Markdown
Contributor

POSIX Compliance (pjdfstest)

============================================================
  pjdfstest POSIX Compliance Results
------------------------------------------------------------
  Files: 130/130 passed    Tests: 832 total (0 subtests failed)
  Result: PASS
------------------------------------------------------------
  Category               Passed    Total   Status
  -------------------- -------- -------- --------
  chflags                     5        5       OK
  chmod                       8        8       OK
  chown                       6        6       OK
  ftruncate                  13       13       OK
  granular                    5        5       OK
  mkdir                       9        9       OK
  open                       19       19       OK
  posix_fallocate             1        1       OK
  rename                     10       10       OK
  rmdir                      11       11       OK
  symlink                    10       10       OK
  truncate                   13       13       OK
  unlink                     11       11       OK
  utimensat                   9        9       OK
============================================================

@github-actions

Copy link
Copy Markdown
Contributor

Benchmark Results

============================================================
  Benchmark — 50MB
------------------------------------------------------------
  Metric                                 FUSE          NFS
  ------------------------------ ------------ ------------
  Sequential read                    169.9 MB/s     200.7 MB/s
  Sequential re-read                2082.0 MB/s    1999.8 MB/s
  Range read (1MB@25MB)                0.3 ms         0.2 ms
  Random reads (100x4KB avg)           0.0 ms         0.0 ms
  Sequential write (FUSE)           1314.1 MB/s
  Close latency (CAS+Hub)            0.121 s
  Write end-to-end                   314.3 MB/s
  Dedup write                       1471.2 MB/s
  Dedup close latency                0.098 s
  Dedup end-to-end                   378.0 MB/s
============================================================
============================================================
  Benchmark — 200MB
------------------------------------------------------------
  Metric                                 FUSE          NFS
  ------------------------------ ------------ ------------
  Sequential read                    545.1 MB/s     767.6 MB/s
  Sequential re-read                2090.7 MB/s    2219.5 MB/s
  Range read (1MB@25MB)                0.2 ms         0.2 ms
  Random reads (100x4KB avg)           0.0 ms         0.0 ms
  Sequential write (FUSE)           1424.8 MB/s
  Close latency (CAS+Hub)            0.112 s
  Write end-to-end                   793.6 MB/s
  Dedup write                       1416.3 MB/s
  Dedup close latency                0.420 s
  Dedup end-to-end                   356.3 MB/s
============================================================
============================================================
  Benchmark — 500MB
------------------------------------------------------------
  Metric                                 FUSE          NFS
  ------------------------------ ------------ ------------
  Sequential read                   1129.2 MB/s    1152.2 MB/s
  Sequential re-read                2101.1 MB/s    2210.3 MB/s
  Range read (1MB@25MB)                0.2 ms         0.2 ms
  Random reads (100x4KB avg)           0.0 ms         0.0 ms
  Sequential write (FUSE)           1387.3 MB/s
  Close latency (CAS+Hub)            0.118 s
  Write end-to-end                  1045.7 MB/s
  Dedup write                       1405.7 MB/s
  Dedup close latency                0.138 s
  Dedup end-to-end                  1011.8 MB/s
============================================================
============================================================
  fio Benchmark Results
------------------------------------------------------------
  Job                        FUSE MB/s   NFS MB/s  FUSE IOPS   NFS IOPS
  ------------------------- ---------- ---------- ---------- ----------
  seq-read-100M                  431.0      465.1                      
  seq-reread-100M               2941.2        7.1                      
  rand-read-4k-100M                0.1        0.1         17         20
  seq-read-5x10M                 641.0      793.7                      
  rand-read-10x1M                  0.1        0.2         36         39
  Random Read Latency           FUSE avg      NFS avg
  ------------------------- ------------ ------------
  rand-read-4k-100M           59897.9 us   50010.1 us
  rand-read-10x1M             27811.8 us   25664.2 us
============================================================

XciD added a commit that referenced this pull request Sep 30, 2026
Fold in the parts of #241 that still apply with RemoteReader: a download
buffer limit of 1 GiB (one fast reader alone keeps up to 256 MiB in
flight, the whole former limit), a download concurrency cap of 124, and
its regression test for #234 (many cold readers, 4 FUSE workers, 5 s
fetch timeout), which now looks for the RemoteReader timeout message.

#241's per-stream download buffers are not needed: fetches are bounded
and drained as they arrive, so no stream holds buffer while its reader
waits for a FUSE worker.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

prefetch hangs / deadlocks in some way during mass fetch

1 participant