Skip to content

Perf: lower init peak - #130

Open
loiccoyle wants to merge 3 commits into
tldr-group:mainfrom
loiccoyle:perf/lower-init-peak
Open

loiccoyle wants to merge 3 commits into
tldr-group:mainfrom
loiccoyle:perf/lower-init-peak

Conversation

@loiccoyle

Copy link
Copy Markdown
Contributor

I did a.second pass of cursoring memory gains, I leave this up to you if you think it is worth merging, it is not required for anything I'm working on. Feel free to close this PR if there is nothing of value :). Thanks for reviewing and getting the previous one in!

Second performance pass to get 1024 cubes through:

  • Sweep the SOR update in x slabs instead of over the whole volume. Every neighbour of an active chequerboard voxel has the opposite colour, so it is never written during the same iteration and slab-wise in-place updating is exact. Bounds the increment buffer to a fixed budget instead of 4 bytes/voxel.
  • Drop the two full-volume chequerboard masks and select the active colour through strided views (-2 bytes/voxel).
  • Reduce ElectrodeSolver.compute_metrics from a full flux field to a chunked (batch, Nx) reduction, which was the largest transient at 8 bytes/voxel.
  • Keep the conductive mask boolean and the reactive neighbour counts uint8 rather than widening both to float32.
  • Compute neighbour sums directly from the unpadded volume with explicit ghost values, removing the padded copy each init made.
  • Write the initial field straight into the padded buffer instead of building mask * profile separately.
  • Release cached blocks around the largest init allocation; the reserved pool for a 1024^3 solve drops from 20.3 to 14.7 GiB.

Tested to be equivalent to the current implementation, and no speed decrease.

loiccoyle and others added 3 commits July 30, 2026 15:25
Large volumes (1024^3) could not be initialised on a 23 GB card: peak
usage
was 22 bytes per voxel, so the buffers alone needed ~25 GB. The
reductions,
all bit-for-bit equivalent to the previous results:

- Sweep the SOR update in x slabs instead of over the whole volume.
Every
  neighbour of an active chequerboard voxel has the opposite colour, so
it is
  never written during the same iteration and slab-wise in-place
updating is
  exact. Bounds the increment buffer to a fixed budget instead of 4
bytes/voxel.
- Drop the two full-volume chequerboard masks and select the active
colour
  through strided views (-2 bytes/voxel).
- Reduce ElectrodeSolver.compute_metrics from a full flux field to a
chunked
  (batch, Nx) reduction, which was the largest transient at 8
bytes/voxel.
- Keep the conductive mask boolean and the reactive neighbour counts
uint8
  rather than widening both to float32.
- Compute neighbour sums directly from the unpadded volume with explicit
ghost
  values, removing the padded copy each init made.
- Write the initial field straight into the padded buffer instead of
building
  mask * profile separately.
- Release cached blocks around the largest init allocation; the reserved
pool
  for a 1024^3 solve drops from 20.3 to 14.7 GiB.

Throughput is unchanged (0.1084 s/iter at 768^3, before and after).
At 256 MiB the chunk covered the whole volume below roughly 512^3, so
mid-sized
images kept full-volume scratch buffers and peaked at 17 bytes/voxel
instead of
10. Dropping to 64 MiB makes the peak a uniform ~10 bytes/voxel across
sizes for
under 1% throughput cost (768^3: 0.1088 vs 0.1084 s/iter).
The reduction hardcoded float32, so a solver constructed with
precision=torch.double silently produced single-precision volume fractions
where it previously returned float64. Accumulate in self.precision instead.

Co-authored-by: Cursor <cursoragent@cursor.com>
@loiccoyle

Copy link
Copy Markdown
Contributor Author

FYI: this is what I'm getting on the benchmark
image
image

@daubners

Copy link
Copy Markdown
Contributor

Very interesting, thanks! Definitely makes sense that exploiting the structure of the SOR scheme saves a lot of memory (almost -20%). I'll have a more detailed look but my current gut feeling is that we would trade a lot of readability and extensibility (from a human perspective) for performance but in a research code the first two are definitely important :D

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants