Repository navigation
Conversation
Every read stream competed FIFO for xet-core's single global download buffer, which hf-mount capped at 256 MiB. Buffer held by downloaded but unconsumed data is only released when a FUSE worker thread consumes it, and under many concurrent cold readers those threads were all blocked waiting on streams that could not get their first term. Reads timed out at cursor=0, retried at the back of the queue, and ended in EIO. Each stream now gets a private AdjustableSemaphore via FileReconstructor::with_buffer_semaphore, sized as a fair share of the download buffer limit (1 GiB by default) clamped to [one xorb, 256 MiB]. The floor cannot go below one xorb: xet-core clamps a term acquire only when it is issued, so a shrink below an in-flight acquire would leave it unsatisfiable. Measured on c7gn.16xlarge, 128 x 1 GiB files, parallel -j64 sha512sum -c: main 996 s at 1.2 Gbps with 108 stream timeouts; fix 59 s at 19.3 Gbps with none. Single-stream reads go from 0.8-1.1 Gbps to 1.2-1.35 Gbps. Adds test_fuse_parallel_cold_reads (24 cold readers on 4 FUSE threads), which fails on main with 13 of 24 readers in EIO.
With per-stream buffers, many concurrent readers keep more terms in flight than xet-core's default cap of 64 connections allows. 124 is the value of xet-core's own high-performance preset. Not measured in isolation; the #234 bench numbers were taken with this cap.
Contributor
POSIX Compliance (pjdfstest) |
Contributor
Benchmark Results |
XciD
added a commit
that referenced
this pull request
Sep 30, 2026
Fold in the parts of #241 that still apply with RemoteReader: a download buffer limit of 1 GiB (one fast reader alone keeps up to 256 MiB in flight, the whole former limit), a download concurrency cap of 124, and its regression test for #234 (many cold readers, 4 FUSE workers, 5 s fetch timeout), which now looks for the RemoteReader timeout message. #241's per-stream download buffers are not needed: fetches are bounded and drained as they arrive, so no stream holds buffer while its reader waits for a FUSE worker.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #234.
Problem
Under many concurrent cold sequential readers (
parallel -j64 sha512sum -cover the mount), reads timed out atcursor=0, retried, and ended inInput/output error. Throughput collapsed to about 1 Gbps on a 200 Gbps link.Every read stream competed FIFO for xet-core's single global download buffer, which hf-mount capped at 256 MiB. Buffer held by downloaded but unconsumed data is only released when a FUSE worker thread consumes it, and those threads were all blocked waiting on streams that could not get buffer for their first term. The 30 s fetch timeout broke the cycle, but the retry re-entered the FIFO at the back, so it mostly re-downloaded and eventually gave up with EIO.
--direct-ioonly reduced kernel readahead piling onto the same handle; the starvation was the same.Fix
Each stream gets a private
AdjustableSemaphorethroughFileReconstructor::with_buffer_semaphore, so streams never wait on each other. The size is a fair share ofHF_XET_RECONSTRUCTION_DOWNLOAD_BUFFER_LIMIT(now 1 GiB by default) clamped to[one xorb, 256 MiB], rebalanced when a stream opens or closes. A lone reader keeps deep pipelining; many readers each keep a guaranteed floor. In-flight memory is bounded bymax(limit, streams * 64 MiB).The floor cannot go below one xorb (64 MiB): xet-core clamps a term acquire against the semaphore total only when it is issued and does not re-clamp pending acquires on shrink, so a smaller floor left in-flight acquires that could never complete (seen as simultaneous timeouts at
cursor=24 MiBduring development).A second commit raises the adaptive download concurrency cap from 64 to 124 (xet-core's high-performance preset value). It is not measured in isolation; the numbers below include it.
Measurements
c7gn.16xlarge (64 vCPU, 200 Gbps, arm64, kernel 6.8), 128 x 1 GiB files, the reporter's command with
--max-threads 64 --no-disk-cache, cold cache:--direct-ioSingle sequential reader, cold: 0.8 to 1.1 Gbps on main, 1.2 to 1.35 Gbps with the fix. Daemon peak RSS under the 64-reader run: 7.8 GiB on main, 8.9 GiB with the fix (the per-handle readahead windows dominate in both).
Tests
test_fuse_parallel_cold_reads: 24 cold readers on 4 FUSE threads with a 5 s fetch timeout. Fails on main with 13 of 24 readers in EIO and 84 stream timeouts; passes with the fix in about 5 s.fuse_opssuite green on the Linux instance; 420 unit tests green.