Release netCDF memory-map pages after each frame - #5463
Conversation
NCDFReader._read_frame now drops the pages of scipy's memory map once the frame has been copied into the Timestep. Nothing else unmaps them, so without this the resident memory of the process grows towards the size of the whole trajectory file while it is read. The map is read only, so dropping is safe: pages are faulted back in if the data are read again. MADV_DONTNEED does not exist on Windows, where the call is skipped and the reader is unchanged.
Documentation build overview
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## develop #5463 +/- ##
========================================
Coverage 93.87% 93.88%
========================================
Files 182 182
Lines 22522 22526 +4
Branches 3206 3207 +1
========================================
+ Hits 21143 21148 +5
Misses 917 917
+ Partials 462 461 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
| The pages of the memory map are released after every frame, so that | ||
| memory use no longer grows towards the size of the trajectory file | ||
| while it is read. Requires ``MADV_DONTNEED``, which is not available | ||
| on Windows. |
There was a problem hiding this comment.
FWIW, I don't think it is necessary to add a versionchanged directive for a bug fix--mostly just appropriate for user-facing behavior changes, otherwise we may have way too many such directives.
| SimpleNamespace(madvise=advice.append), | ||
| ) | ||
| universe.trajectory[1] | ||
| assert advice == [mmap.MADV_DONTNEED] |
There was a problem hiding this comment.
The test is a bit convoluted with monkeypatch--could we avoid that?
Also, am I understanding correctly that on this branch something like https://github.com/pythonprofilers/memory_profiler should show a constant level of physical memory consumption, while it should increase linearly when reading in/iterating over a trajecotry of the appropriate type on the develop branch?
We could probably capture that with https://asv.readthedocs.io/en/stable/writing_benchmarks.html#peak-memory as well, assuming you're saying that physical memory is being used excessively on develop in these scenarios.
Description
NCDFReader._read_frame now drops the pages of scipy's memory map once the frame has been copied into the Timestep. Nothing else unmaps them, so without this the resident memory of the process grows towards the size of the whole trajectory file while it is read, as documented in discussion here. (ping @orbeckst )
MADV_DONTNEED does not exist on Windows, so the call is guarded by hasattr(mmap, "MADV_DONTNEED") and the reader is unchanged there.
Changes made in this Pull Request:
NCDFReader._read_framenow drops the pages of scipy's memory map once theframe has been read into the
Timestep.LLM / AI generated code disclosure
LLMs or other AI-powered tools (beyond simple IDE use cases) were used in this contribution: yes test written by AI
PR Checklist
package/CHANGELOGfile updated?package/AUTHORS?Developers Certificate of Origin
I certify that I can submit this code contribution as described in the Developer Certificate of Origin, under the MDAnalysis LICENSE.