Feat/multi merkle tree gpu resident - #951
Draft
ColoCarletti wants to merge 26 commits into
Draft
ColoCarletti wants to merge 26 commits into
ColoCarletti wants to merge 26 commits into
Conversation
…he batched prover
…r (hybrid resident handles)
…E download for plain tables)
Per table the batched prover still ran a redundant HOST LDE FFT (main+aux) to feed prep trees / retained_main / later phases, leaving the GPU idle. This completes the device port: - aux/main: resident device expand first; host-expand only for the small-table fallback / Retain. A per-table commit-stream sync closes an async race the old host FFT delay had hidden. - OOD (phase 4): device-recompute device-only + read the trace OOD via the GPU barycentric fast path (no host coset eval, no D2H download). Byte-identical proof (roots unchanged); all batched tests pass. e22 mainnet prove 284s -> 155s (-45%), GPU util 17% -> 44%. Budget test updated: phase 4 is now a full device expansion.
Route the batched aux build through the resident GPU build (fingerprints + term columns + running-sum accumulate) and bulk-download the result into the host aux table, and transpose the main trace to column-major on device (ResidentMain::HostRowMajor) instead of on the host. Removes the host set_aux writes, the host accumulate and the ~1.36s/epoch columns_main transpose. Byte-identical; aux_commit 3.64s -> 1.52s/epoch.
Upload the composition parts to a device handle in the batched deep_codeword so the DEEP composition takes the fully-resident arm (device parts + device inv-denoms) instead of the host build_r4_inv_denoms_cpu batch-inverse. Byte-identical; deep_fri 3.96s -> 2.22s/epoch.
The parts D2H de-interleaved the 3 ext3 slabs into per-column row-major ext3 on the host (a strided gather, ~40% of the download). Do it on device via a new interleave_ext3_slabs kernel + per-part D2H straight into the owning buffers (no host copy). Byte-identical; comp_parts 3.02s -> 2.63s/epoch.
… ext3 downloads Move the composition-parts D2H de-interleave onto the GPU (new interleave_ext3_slabs kernel + per-part D2H into the owning buffers), and reinterpret the aux-trace and DEEP-codeword ext3 downloads in place instead of the per-element u64_to_ext3_vec host copy. Byte-identical; comp_parts 3.02->2.63s, aux_commit 1.52->1.29s, deep_fri 2.22->2.0s per epoch.
jotabulacios
added a commit
that referenced
this pull request
Sep 18, 2026
The batched proof format does not need building: #951 has it, with one mixed-height MMCS per round and ONE FRI instance per epoch — which is what the spec asks for, rather than the group-by-exact-height fallback this branch was heading towards. It was deleted from another lane by #973 and #974 and is preserved at archive/batched-format-pre-deletion. So this brings the pieces over rather than reinventing them. What comes across is everything the format needs except the driver: crypto/crypto merkle_tree traits + the field-element backend they need crypto/stark fri/mmcs.rs the mixed-height tree fri/batched.rs height combination, batched commit phase, shared challenge derivation batched/round4.rs the round-4 transcript sequence par.rs par_for_each_mut_indexed, which #974 deleted 24 of their tests come with it and pass, including streaming_builder_serves_the_base_group_without_holding_it — the property this branch needs, since a builder that absorbs each table's LDE and frees it is what lets the batched rounds run without holding all of them. What is deliberately NOT brought is #951's own driver. It takes every table's trace at once, which is the residency this branch exists to remove: its floor is all 227 traces resident, measured at 41077 MB on the ethrex block, and this branch's whole proof fits in 21952 MB. The two remove different things — #951 the simultaneous LDEs, this the simultaneous traces — so the driver is the piece to write rather than to copy, and its header is the specification for it. round4's own five tests are not collected yet; they lean on parts of the format that are not across.
jotabulacios
added a commit
that referenced
this pull request
Sep 18, 2026
Brought in by 6917f6b and reverted unchanged: it is the next step, not this one. The spec's Approach 1 asks for a batched FRI — "accumulate FRI polys into one batch polynomial" — and that is deep_for_table plus batch_fri, which are already here and already checked against a real proof. Grouping their output by exact height gives 13 FRI instances where there are 227, which is the ~54% of proof size the histogram priced. The mixed-height MMCS is a different thing: it batches the COMMITMENTS, one tree per round instead of one per table, and takes 13 FRI instances down to 1. It is strictly more, and it costs a new proof format, a new verifier and the recursion ELFs behind them. Worth doing, after. Nothing is lost by taking it out: it is in this history, and in #951 and archive/batched-format-pre-deletion. Reverting the revert brings it back. And none of it bears on memory, which is what Approach 1 was for and which is already done — 21952 MB against main's 110261 MB, proof verified.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.