Defer the data-parallel gradient all-reduce to update() under gradient accumulation - #5099
Open
NuojCheng wants to merge 1 commit into
Open
Defer the data-parallel gradient all-reduce to update() under gradient accumulation#5099NuojCheng wants to merge 1 commit into
NuojCheng wants to merge 1 commit into
Google CLA / cla/google
succeeded
Sep 2, 2026 in 8s
✅ All contributors are covered under a CLA with Google
See https://cla.developers.google.com/ for more info about Google's Contributor License Agreement (CLA).
ℹ️ Googlers: Go here to view more details and manage scans for this pull request.
Details
The following contributors were found for this pull request:
✅ f3e4046 Author: @NuojCheng <che******in@google.com>
(Only the first commit for a unique contributor is listed.)
Loading