Skip to content

Reduce memory and intermediate work in packing and minimization - #587

Merged
kierandidi merged 9 commits into
masterfrom
codex/pack-memory-20261010
Oct 11, 2026
Merged

kierandidi merged 9 commits into
masterfrom
codex/pack-memory-20261010

Conversation

@kierandidi

@kierandidi kierandidi commented Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator

Scoring, repacking, and minimization retain large temporary arrays that can be avoided without changing public APIs or numerical behavior. This combines the performance work from #587, #588 and #589:

  • Share the CUDA annealer's sequential quench workspaces; fix empty-pose boundary reads exposed by the expanded regression coverage.
  • Construct rotamer topology from real atoms and transfer six forest fields together, avoiding padded host intermediates and repeated transfers.
  • Keep circular L-BFGS histories in place and reorder only small matrices/vectors. Make packing benchmark samples reset inputs/RNG outside timing and finish CUDA work before returning.

Matched measurements: CUDA workspace peak falls by 10.10 MiB in the capacity probe; CPU rotamer construction improves from 120.25 to 111.35 ms (one pose) and 337.01 to 310.95 ms (three poses); ragged forest host peak falls from 22.87 to 10.99 MB. L-BFGS helpers improve 1.09–1.23× at histories 32/128, with operator allocation totals falling from 36.9 to 0.80 MB in the large-history probe. History 8 can be about 3% slower (~2 μs). These are component results: full packing/FastRelax timing did not establish an end-to-end speedup. Separate-process packing results were noisy/slower; alternating same-process comparisons were neutral.

Validation: 405 CPU packing tests and 120 CPU optimization tests passed, plus independent recurrence/autograd and real-minimization parity checks. On L40S / CUDA 13 / PyTorch 2.10.0a0, 136 integrated CUDA tests passed; the final empty-pose fix then passed all 31 streaming-graph CUDA tests and 15 focused CPU tests. Seeded comparisons preserve scores, coordinates, assignments, and subsequent RNG draws. The consolidated components exactly match those tested commits; formatting and diff checks pass. GitHub CI and repository review remain the merge gates.

Companion PRs: #583 for environment/release/docs; #590 for cleanup. This PR absorbs #588 and #589.

@kierandidi
kierandidi marked this pull request as ready for review October 11, 2026 04:34
@kierandidi kierandidi changed the title Reuse CUDA annealer quench permutation workspace Reduce memory and intermediate work in packing and minimization Oct 11, 2026
@kierandidi
kierandidi merged commit 2a306ad into master Oct 11, 2026
3 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant