Skip to content

Feat/multi merkle tree gpu resident - #951

Draft
ColoCarletti wants to merge 26 commits into
mainfrom
feat/multi-merkle-tree-gpu-resident
Draft

ColoCarletti wants to merge 26 commits into
mainfrom
feat/multi-merkle-tree-gpu-resident

Conversation

@ColoCarletti

Copy link
Copy Markdown
Collaborator

No description provided.

Per table the batched prover still ran a redundant HOST LDE FFT (main+aux)
to feed prep trees / retained_main / later phases, leaving the GPU idle.
This completes the device port:
- aux/main: resident device expand first; host-expand only for the
  small-table fallback / Retain. A per-table commit-stream sync closes an
  async race the old host FFT delay had hidden.
- OOD (phase 4): device-recompute device-only + read the trace OOD via the
  GPU barycentric fast path (no host coset eval, no D2H download).

Byte-identical proof (roots unchanged); all batched tests pass. e22 mainnet
prove 284s -> 155s (-45%), GPU util 17% -> 44%. Budget test updated: phase 4
is now a full device expansion.
Route the batched aux build through the resident GPU build (fingerprints +
term columns + running-sum accumulate) and bulk-download the result into the
host aux table, and transpose the main trace to column-major on device
(ResidentMain::HostRowMajor) instead of on the host. Removes the host
set_aux writes, the host accumulate and the ~1.36s/epoch columns_main
transpose. Byte-identical; aux_commit 3.64s -> 1.52s/epoch.
Upload the composition parts to a device handle in the batched deep_codeword
so the DEEP composition takes the fully-resident arm (device parts + device
inv-denoms) instead of the host build_r4_inv_denoms_cpu batch-inverse.
Byte-identical; deep_fri 3.96s -> 2.22s/epoch.
The parts D2H de-interleaved the 3 ext3 slabs into per-column row-major ext3
on the host (a strided gather, ~40% of the download). Do it on device via a
new interleave_ext3_slabs kernel + per-part D2H straight into the owning
buffers (no host copy). Byte-identical; comp_parts 3.02s -> 2.63s/epoch.
… ext3 downloads

Move the composition-parts D2H de-interleave onto the GPU (new
interleave_ext3_slabs kernel + per-part D2H into the owning buffers), and
reinterpret the aux-trace and DEEP-codeword ext3 downloads in place instead of
the per-element u64_to_ext3_vec host copy. Byte-identical; comp_parts
3.02->2.63s, aux_commit 1.52->1.29s, deep_fri 2.22->2.0s per epoch.
jotabulacios added a commit that referenced this pull request Sep 18, 2026
The batched proof format does not need building: #951 has it, with one
mixed-height MMCS per round and ONE FRI instance per epoch — which is what the
spec asks for, rather than the group-by-exact-height fallback this branch was
heading towards. It was deleted from another lane by #973 and #974 and is
preserved at archive/batched-format-pre-deletion.

So this brings the pieces over rather than reinventing them. What comes across
is everything the format needs except the driver:

    crypto/crypto  merkle_tree traits + the field-element backend they need
    crypto/stark   fri/mmcs.rs      the mixed-height tree
                   fri/batched.rs   height combination, batched commit phase,
                                    shared challenge derivation
                   batched/round4.rs  the round-4 transcript sequence
                   par.rs           par_for_each_mut_indexed, which #974 deleted

24 of their tests come with it and pass, including
streaming_builder_serves_the_base_group_without_holding_it — the property this
branch needs, since a builder that absorbs each table's LDE and frees it is what
lets the batched rounds run without holding all of them.

What is deliberately NOT brought is #951's own driver. It takes every table's
trace at once, which is the residency this branch exists to remove: its floor is
all 227 traces resident, measured at 41077 MB on the ethrex block, and this
branch's whole proof fits in 21952 MB. The two remove different things — #951
the simultaneous LDEs, this the simultaneous traces — so the driver is the piece
to write rather than to copy, and its header is the specification for it.

round4's own five tests are not collected yet; they lean on parts of the format
that are not across.
jotabulacios added a commit that referenced this pull request Sep 18, 2026
Brought in by 6917f6b and reverted unchanged: it is the next step, not this
one. The spec's Approach 1 asks for a batched FRI — "accumulate FRI polys into
one batch polynomial" — and that is deep_for_table plus batch_fri, which are
already here and already checked against a real proof. Grouping their output by
exact height gives 13 FRI instances where there are 227, which is the ~54% of
proof size the histogram priced.

The mixed-height MMCS is a different thing: it batches the COMMITMENTS, one tree
per round instead of one per table, and takes 13 FRI instances down to 1. It is
strictly more, and it costs a new proof format, a new verifier and the recursion
ELFs behind them. Worth doing, after.

Nothing is lost by taking it out: it is in this history, and in #951 and
archive/batched-format-pre-deletion. Reverting the revert brings it back.

And none of it bears on memory, which is what Approach 1 was for and which is
already done — 21952 MB against main's 110261 MB, proof verified.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant