Skip to content

bcachefs: ec_read_around: widen on failure (tools#1116 follow-up) - #141

Open
aderk wants to merge 6 commits into
koverstreet:masterfrom
aderk:ec-read-around-widen
Open

aderk wants to merge 6 commits into
koverstreet:masterfrom
aderk:ec-read-around-widen

Conversation

@aderk

@aderk aderk commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Tests for the widen-on-failure follow-up to koverstreet/bcachefs-tools#1116, on top of #136: ec_read_around_bad_block_no_slow_wait and ec_repair_bad_block_no_slow_read (fail on #1116's tip, pass with the follow-up), and ec_read_around_fallback_reconstruct now reads around device 1.

aderk added 5 commits October 5, 2026 19:06
Tests for reading around a member much slower than the rest of its
stripe: a read of its data is reconstructed from the stripe's other
blocks when its latency estimate says that's faster. Five or six 1G
members, 3+1 stripes unless noted, and a dm-delay member at 500 ms.

- ec_read_slow_device: reads of the slow member's data don't wait for it.
- ec_read_healthy, ec_read_around_disabled: no read-arounds on a healthy
  pool, or with ec_read_around_penalty=0.
- ec_read_around_pq: 4+2 stripes with two slow members.
- ec_read_around_fallback, _fallback_reconstruct: a read-around that
  reads a bad block falls back to the member, then to the full
  reconstruct.
- ec_read_around_small_extents, ec_read_around_nocsum: read
  amplification with 4k extents, and unchecksummed data.
- ec_repair_reads_around_slow_device, ec_reconcile_reads_around_slow_device:
  evacuation and reconcile rebuild a slow member's blocks from the rest
  of the stripe.
- ec_reconcile_learns_failing_device, ec_evacuate_learns_failing_device:
  from a cold estimate, with a member at 7 s per read, reconcile's and
  evacuation's own reads bring its estimate up.
- ec_reconstruct_skips_failed_block: a reconstruct after a read error
  doesn't read the failed block again.

Tests with a delayed member warm its estimate first with random reads,
read-around off. Needs the read-around series. Replaces koverstreet#130, whose
tests asserted that an evacuating device is never read.
With a promote target, a read-around promotes as a direct read does, so
a second read of a slow member's data comes from the promote target.
Five members in group hdd, two behind dm-delay, and a promote target in
group ssd, set only after the setup's reads and the prewarm. The first
pass reads around; the second must not read around at all, and must
read most of the file from the promote target.

ec_ra_setup takes per-device labels from ec_ra_labels.

Needs the read-around series with "ec: promote read-around results".
4+2 stripes with a slow member whose own block is good and a parity
block on a fast member bad: the read-around reads the bad block and
fails. Reading the one block it left out gives one block more than a
rebuild needs, so the data can be rebuilt around the bad block without
waiting on the slow member.

The bad block is parity so that no read of the file's own data fails
and goes to the full reconstruct, which reads every block. Reads over
200 ms are allowed up to ec_ra_coin_reads, as the read path still sends
a read to the slow member with probability (r/d)^2.

Fails until the read-around widens on failure.
Evacuating the slow member of 4+2 stripes with a bad parity block in 16
of them: repair reads around the two slowest members of each stripe,
finds the bad block among the four it reads, and reads the faster of
the two it left out to rebuild the slow member's block. Nothing may be
read from the slow member's blocks; repair has no coin toss.

The bad block is parity because a bad data block's data is moved, and
the move's read of it goes to the full reconstruct, which reads every
block, the slow member's included. Seven members, so each stripe is
rebuilt in place: with six, every stripe is narrowed and its data moved,
and the moves' transaction restarts at read submit are counted as
data_update_fail, over the end checks' limit with or without this
change.

Fails until repair widens by one block on failure.
With a failed read-around widened by the block it left out, a
read-around of device 0 with device 0's block and one other bad
rebuilds around the other, and the direct read of device 0 that the
test is about no longer happens; the read-arounds that still failed
were those of device 1's data, if device 1 held data in that stripe.

Make device 1's block and another bad, in a stripe with data on device
1. Device 0, the block left out of a read-around of device 1, is
slower than device 1, so the read-around isn't widened: it fails, the
direct read gets device 1's bad block, and the full reconstruct has two
bad blocks, as before.
The seventh member was there only because, with six, the moves'
transaction restarts at read submit were counted as data_update_fail,
over the end checks' limit. The read-around pick no longer relocks a
data update's transaction (tools#1116), so drop it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant