Share chunk-memory sizing - #1013
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/structured-data-models/.coderabbit.yaml Review profile: QUIET Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughAutomatic chunk sizing now uses a shared CUDA memory-budget helper. Transformer attention delegates batch-size calculation to a method that also handles non-CUDA devices. ChangesAutomatic chunk sizing
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Merge Risk: ⚪ Minimal · up to The shared chunk-sizing calculation appears to preserve existing behavior. No actionable merge-blocking risk remains after normal checks. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
96cb79a to
219ebe7
Compare
- Move the `SDM_CHUNK_MEMORY_FRACTION` budget of attention and TabFM cell embedding chunks into one helper, `sdm._memory.chunk_memory_limit`. - Expose the automatic attention batch size limit as `TransformerBlock.auto_batch_size_limit`, so callers can plan passes that align with its chunks. Signed-off-by: Jingang Qu <jqu@nvidia.com>
219ebe7 to
62280d8
Compare
Split from #996 to keep each review focused on one behavior or optimization.