This repository runs benchmarks and reports their measurements: build and cache times, storage bytes, cache hits, output checks, failures, and source/run links. It uses one case format, shared execution workflows, and shared reporting. Performance explanations and recommendations belong in a separate review.
Workloads live in cases/. Each case pins its upstream source and declares
its recipe, comparison, cache scope, and output verification. Exploratory technical
work is an evaluation. Promotion changes metadata and suite membership; it does
not change the case's identity or copy its executor.
Follow docs/process.md to add or run a case.
Tool coverage lists configured families and qualification gaps.
AGENTS.md applies the same requirements to agents and humans.
bin/bench catalog generates data/latest/series.json,
the common index for planned, requested, incomplete, failed, and completed
evaluations. It uses the canonical report validator and calculations, retains
evidence gaps, and leaves publication review explicit.
Standard comparisons use two shared fresh and rolling workflows, which also
support workflow_call. All cases use one BoringCache wrapper
for the product invocation and release pin. Preparation and common tool setup
use shared actions; case payloads keep their upstream recipe and output checks.
Docker cases also share provider setup, timing, and publication policy.
bin/bench collect imports preserved phase artifacts and generates a report after
checking the dispatch and completed jobs. Original product evidence is retained
after product cleanup.
Each fresh sample runs independently with its own cache scope. Providers start in parallel; the sample's warm builds follow its cold builds on fresh runners. Other samples and cases do not wait for it. Native rolling comparisons keep a queue for the case and series whose cache they advance; the Cargo chains also keep their seed ordering. GitHub runner availability can still cause waiting, and the repository's Actions Cache quota is shared.
flowchart LR
Case[Case definition] --> Workflow[Shared experiment workflow]
Workflow --> Setup[Shared preparation and tool setup]
Setup --> Provider[BoringCache wrapper or declared comparator]
Provider --> Recipe[Case build and output checks]
Recipe --> Report[Shared records, reports and evidence]
bundle install
bin/bench list
bin/bench check
bin/bench plan hugo-go --lane fresh
bin/bench prepare hugo-go --native-lane fresh --directory /tmp/hugo-benchmark
bin/bench start hugo-go --series screening-01 --lane fresh --samples 2All BoringCache plans use boringcache/benchmarks. Case tags and series identities
control reuse within that workspace. GitHub Actions uses OIDC; connection to the
existing workspace must be verified before cutover.
Compare the declared build and cache reuse operation. Record storage with its provider and measurement source. Queue delay, unrelated dependency setup, and job duration are recorded separately from the measured operation. Keep cold, identical-source warm, and changed-source results separate. Collect the declared samples and report medians and ranges without discarding slow observations.
The product benchmark page uses separately reviewed results. Published interfaces remain available:
data/latest/report.md: latest published cohort reportdata/latest/index.json: machine-readable workload indexdata/latest/providers.json: provider comparisonssuites/published.json: shared publication registryresults/historical-website/report.md: reviewed historical website observations and preserved evidence
Historical run URLs retain their original execution repository. A moved case does not move an Actions run. Evidence preservation exports each attempt, jobs, logs, artifacts, commit, and workflow, records checksums, and reports missing material.
Consolidation is in progress. docs/migration.md lists its
acceptance conditions. Old repositories remain until their execution, callers,
published references, and required evidence are verified. Forks are deferred in
migration/forks.json, outside the active suite. Select
and migrate a fork evaluation when needed.