Conversation
Large volumes (1024^3) could not be initialised on a 23 GB card: peak usage was 22 bytes per voxel, so the buffers alone needed ~25 GB. The reductions, all bit-for-bit equivalent to the previous results: - Sweep the SOR update in x slabs instead of over the whole volume. Every neighbour of an active chequerboard voxel has the opposite colour, so it is never written during the same iteration and slab-wise in-place updating is exact. Bounds the increment buffer to a fixed budget instead of 4 bytes/voxel. - Drop the two full-volume chequerboard masks and select the active colour through strided views (-2 bytes/voxel). - Reduce ElectrodeSolver.compute_metrics from a full flux field to a chunked (batch, Nx) reduction, which was the largest transient at 8 bytes/voxel. - Keep the conductive mask boolean and the reactive neighbour counts uint8 rather than widening both to float32. - Compute neighbour sums directly from the unpadded volume with explicit ghost values, removing the padded copy each init made. - Write the initial field straight into the padded buffer instead of building mask * profile separately. - Release cached blocks around the largest init allocation; the reserved pool for a 1024^3 solve drops from 20.3 to 14.7 GiB. Throughput is unchanged (0.1084 s/iter at 768^3, before and after).
At 256 MiB the chunk covered the whole volume below roughly 512^3, so mid-sized images kept full-volume scratch buffers and peaked at 17 bytes/voxel instead of 10. Dropping to 64 MiB makes the peak a uniform ~10 bytes/voxel across sizes for under 1% throughput cost (768^3: 0.1088 vs 0.1084 s/iter).
The reduction hardcoded float32, so a solver constructed with precision=torch.double silently produced single-precision volume fractions where it previously returned float64. Accumulate in self.precision instead. Co-authored-by: Cursor <cursoragent@cursor.com>
Contributor
Author
Contributor
|
Very interesting, thanks! Definitely makes sense that exploiting the structure of the SOR scheme saves a lot of memory (almost -20%). I'll have a more detailed look but my current gut feeling is that we would trade a lot of readability and extensibility (from a human perspective) for performance but in a research code the first two are definitely important :D |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


I did a.second pass of cursoring memory gains, I leave this up to you if you think it is worth merging, it is not required for anything I'm working on. Feel free to close this PR if there is nothing of value :). Thanks for reviewing and getting the previous one in!
Second performance pass to get 1024 cubes through:
Tested to be equivalent to the current implementation, and no speed decrease.