perf(storer): reuse buffers and eliminate per-chunk allocations in reserve sample - #5612
gacevicljubisa wants to merge 7 commits into
Conversation
Sampling is about to start reading chunks into a buffer that each worker reuses. Nothing today checks that the bytes handed back in a SampleItem are still the bytes of that chunk, so a reused buffer would silently hand the redistribution proof the contents of some later chunk. assertValidSample now checks two things for every item: that ChunkData still reproduces ChunkAddress, and that no two items share a backing array. Every existing sample test picks both up. Also add the rulers for the work that follows. BenchmarkReserveSample1k keeps its name and behaviour so the recorded baseline stays comparable; its body moves to a helper that BenchmarkReserveSample10k reuses over a ten times larger reserve. BenchmarkChunkStoreGet measures a single chunk read, split into a variant that builds the ChunkStore handle per call as the sampler does today and one that hoists it, so the cost of the handle alone is visible. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFUseFQ8rhp6N9X6YS7pbq
db.ChunkStore() builds three objects every time it is called, and the sampler called it once per chunk. On a testnet node that is 2.3 million handles per round, thrown away immediately. Hoist it to one per worker. The handle is deliberately not shared across workers: the read-only chunk store makes no thread-safety promise, and three allocations per worker is already nothing next to three per chunk. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFUseFQ8rhp6N9X6YS7pbq
444d89d to
841bbaf
Compare
Both ReserveSample benchmarks fill their reserve with chunk.GenerateValidRandomChunkAt, which produces content-addressed chunks only. The SOC branch of transformedAddress is therefore never measured by them, and any work that branch does beyond hashing the wrapped CAC is invisible. Benchmark the function directly, one case per chunk type, through a shim in export_test.go. The shim takes a swarm.Chunk and is held to that signature on purpose so the same benchmark source can be run against branches whose internal transformedAddress is shaped differently; only the shim changes between them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFUseFQ8rhp6N9X6YS7pbq
|
|
||
| type ChunkStore interface { | ||
| Getter | ||
| GetterInto |
There was a problem hiding this comment.
Not sure if we need to add it to the ChunkStore interface. Since we only use ReadOnlyChunkStore for sampling, we can only add it to the interface below. I think we can even avoid adding it to transaction.ChunkStore also.
There was a problem hiding this comment.
Dropping it makes ReadOnlyChunkStore no longer a subset of ChunkStore, which breaks the test storages and the mock storer that return the same value as both. Then even more changes are needed...
| return ch, err | ||
| } | ||
|
|
||
| func (c *chunkStoreTrx) GetInto(ctx context.Context, addr swarm.Address, buf []byte) (n int, err error) { |
There was a problem hiding this comment.
Check comment above. Since we only need ReadOnlyChunkStore to provide this function, we can avoid adding it here.
| continue | ||
| } | ||
|
|
||
| ch, err := phase3ChunkStore.Get(ctx, item.chunkAddress) |
There was a problem hiding this comment.
iirc there's another thing to optimize here: notice that there could be many inserts into the sample, however, the sample is fixed size and, finally it converges onto a certain set of chunks. what this means is that you're inserting an item into the sample without knowing whether it would actually survive into the final sample set. so the next plausible thing to do here is to not Get the chunk here - insert the sample item without the chunk data (but leave all the other fields on the type), and backfill the surviving sample members once the sample is finalized i.e. before returning the sample at L318
There was a problem hiding this comment.
you might want to even do the stamp loading at the same step. in fact at this stage you need not save anything into the sample apart from the address and the transformed address
There was a problem hiding this comment.
I have skiped it intentionally, on 4M chunks, it is expected 200-250 chunks to land in phase 3. This is case on testnet and mainnet. Deferring validation could leave the samle with less then 16 items. I would keep this PR as is, and maybe we can have a follow up. Is that fine by you?
There was a problem hiding this comment.
ok 👍 btw after calin's last PR i am no longer sure about my suggestion to rely on cap instead of len. having seen how the std lib does reads - it indeed usually uses len, not cap. so - my bad
There was a problem hiding this comment.
done with additional test.
Checklist
Description
This PR addresses per-chunk allocations and buffer overhead in
ReserveSample(issue #5174):escapes.
actually qualify for the sample candidate set.
BenchmarkReserveSample10k.
A/B on
bee-light-testnet, two nodes with identical config (no SIMD, verbosity 3, 1500m CPU, storage radius 2,isWarmingUp: false). Sample triggered viaGET /rchash/2/<own-overlay>/<anchor2>; allocations attributed withpprof -diff_basefocused on theReserveSamplesubtree.b6537ea2chunkstore.readChunkbuffertransformedAddress(control)bmt.Write(control)Total allocation churn per sample round: 24.65 GB → 10.30 GB over 2.405M chunks.
Open API Spec Version Changes (if applicable)
Motivation and Context (Optional)
Related Issue (Optional)
Screenshots (if appropriate):
AI Disclosure