Read the upstream workflow and relevant build files. Identify its output, normal source changes, existing cache, and the operation the proposed cache can improve. Check existing cases first. Record recipe changes as patches or explicit projections; distinguish them from faithful upstream execution.
Create a blocked draft through the same interface:
bin/bench new <case-id> --repository <owner/repository> --revision <full-sha> --question 'What are we evaluating?'Edit cases/<case-id>/case.json using schemas/case.schema.json. Start with
kind: evaluation and reporting.publication: evaluation-only. Use JSON, exact
source SHAs, and the workspace boringcache/benchmarks.
For a Go workload using the existing pinned-source fresh/rolling shape:
bin/bench new <case-id> --repository <owner/repository> --revision <full-sha> \
--question 'What are we evaluating?' --shape go --tool-version <exact-go-version>This creates the case definition, .gitmodules, product plan, recipe contract
and scripts/verify-output.rb stub. It selects the shared Go action and native
fresh/rolling workflows. Review the command in both the product plan and recipe
contract, replace the placeholder upstream digest with the reviewed file digest,
and implement the output check. The stub fails until implemented. Remove the
case's blocker after that review, prepare the case and verify its output check,
then add it to the appropriate suite after qualification. Do not copy timers, provider setup,
phase reporting or workflows into the new case.
The shared prepare-native action checks pinned source and recipe files.
native-cache-benchmark owns provider selection, publication policy, timing,
output-check invocation, phase records and artifact upload for Go, Maven,
Gradle, n8n Turbo and Storybook. Dependency preparation scripts run between
the cache-setup timer and build timer. Java cases retain their toolchain setup and comparator configuration;
their output checks live in case-owned Ruby scripts. The Go action combines Go
setup with these shared steps, so another Go case needs no case action.
Declare the question, providers, primary metric, timed scope, storage measurement,
cache scope, sample count, workflows, and output verification. Workflow payloads
live in cases/<case-id>/payload/. Docker workloads use recipe.json and plans/
with the shared Docker adapter. Do not vendor upstream source: bin/bench prepare
obtains the declared revision and applies recorded patches in a disposable checkout.
Record reviewed upstream file digests and build commands in recipe-contract.json.
When a case has distinct pinned recipes, use a named contract with the same
verify-upstream-recipe.rb <source-directory> <contract-file> verifier. A contract
can declare groups of plan paths with a shared adapter and command under plans;
the verifier checks every named plan. A digest mismatch requires reviewing the
source, command, and environment changes before updating the contract. Zed's
recipe review distinguishes its older Cargo layer
proof from its newer rolling recipe.
For the standard BoringCache/Actions Cache comparison, use
native-fresh-benchmark.yml and native-rolling-benchmark.yml. Declare the shared
or case composite action, benchmark ID, toolchain inputs, and lane-specific publication
settings in execution.native. Set case_id in the case's workflow selectors. The
shared executor prepares that reviewed action on the disposable runner; agents
do not copy workflow setup or reporting into each new case. Declare tag mappings
in execution.cache_tags and use scope-case-cache.rb. Use prepare-source.rb
for an unchanged pinned upstream checkout.
When one case has several workloads, declare their recipe overrides in
execution.native.variants and register the same variants on the shared workflow
entry with variant_input: variant. Select one variant before execution:
bin/bench plan n8n --variant turbo
bin/bench prepare n8n --native-lane fresh --variant turbo --directory /tmp/n8n-turbo
bin/bench start n8n --series screening-turbo-01 --variant turbo --samples 1The selected variant fixes the case action, benchmark ID, toolchain inputs, and
report identity. Provider-specific boolean flags belong in provider_flags;
for example, the Actions Cache arm keeps BoringCache's tool-proxy flags disabled.
All variants use the same provider lifecycle and canonical reporter. Missing or
unknown variants fail before dispatch. Contract checks prepare every registered
native variant, including its report filenames. A fresh variant passing does not
qualify another variant or the original rolling matrix.
Native planning also rejects a phase action that does not report verified output.
Add and review the actual output check before setting --verified-output; the
flag alone does not verify a workload. Schema and file-contract checks can pass
while an imported recipe still lacks that execution requirement.
For native Docker comparisons that declare load_image, fresh inputs must set
it to true. Both arms then export and load their built image during the timed
build. The shared verify-docker-output.rb checks the selected provider's image
after timing and passes verification to the reporter only after inspection
succeeds. The image check establishes that the build produced the declared image;
application behavior requires any additional checks declared by the case.
Publication and cache-only projections remain diagnostic until their output and
comparison boundary have been reviewed. Do not add the verification flag to a
cache-only build.
Keep different provider sets, architectures, source-change sequences, and output
behavior explicit in their adapters or case actions. OBS, Cargo product proofs,
and the Docker corpus retain their distinct execution shapes. Add a workflow
only when the experiment requires a different shape, and register it in the case.
Use the canonical Ruby reporter for every execution path. Add
suite membership for an explicit scheduling purpose. suites/published.json is
the shared registry for published reporting.
suites/scheduled.json selects the maintained suite and new native families for
the proposed shared cadence. See benchmark schedules for current
ownership and the remaining source and rolling work.
.github/actions/boringcache is the sole BoringCache One invocation and release
pin. It forwards the product inputs used by the cases and retains raw evidence.
Tool setup uses
shared actions such as prepare-case, setup-docker, and setup-node; case
actions retain their upstream commands, patches, and output verification. Native
fresh and rolling workflows support both dispatch and workflow_call. Use them
for evaluations and published cases without copying the provider lifecycle.
Reusable callers declare contents: read, actions: read, packages: read, and
id-token: write for the called job. GitHub checks these permissions when loading
the workflow, including calls whose execution condition is false.
Docker tool caches remain product-plan settings rather than a second implementation
of the product's proxy or cache lifecycle.
Docker case actions call .github/actions/docker-benchmark for provider setup,
timing, and execution. One phase policy controls cache publication in both arms:
cold publishes; an identical-source warm replay restores without publishing;
an explicitly declared publish-on-warm observation publishes in both arms.
Image loading stays inside this action's timer. Output inspection runs after it.
Use benchmark-phase.rb scope for the shared case/series identity, and keep
tool-specific tags in the product plan.
The provider wrapper registers .github/actions/retain-product-evidence before
calling One. Its GitHub post hook runs after One's cleanup, then uploads the
original final evidence JSON using GitHub's artifact SDK. Missing or invalid
evidence fails retention. The small JavaScript hook is required for GitHub post
steps; it does not interpret product measurements or implement cache behavior.
Its dependencies install during preparation, outside the comparison timer.
The Ruby reporter retains its structured measurements; the original envelope
remains a separate artifact so post-step status and future product fields survive.
Case actions upload their phase records. They do not upload a second copy of
product evidence before cleanup.
Storage measurements pass the resolved tags to boringcache check --no-git --no-platform --json. The CLI checks those exact tags; the former --exact
option is unnecessary and is rejected by CLI 1.34.0. A failed probe leaves storage
unmeasured.
The harness Ruby version is pinned once in .tool-versions. Ruby setup reads
that file from the harness checkout, including when a shared action is called
from an older benchmark repository. Guardrails verify preparation on Ubuntu
26.04 and macOS 26 as well as running the full harness checks on Ubuntu 24.04.
Runtime support must be verified before selecting a new runner or Ruby version.
BoringCache Action cache failures fail benchmark jobs. The wrapper sets that
policy literally and converts the optional cache-miss input to a boolean expression.
Retain native tool error counters separately. Successful outputs do not establish
an error-free native cache operation; review those counters before making a claim.
Nested composite post steps can receive a different input context; forwarding a
boolean as an unchecked string can become empty during cleanup. Qualification
must inspect post-job logs and evidence as well as the job conclusion.
bin/bench check <case-id>
bin/bench plan <case-id> --lane fresh
bin/bench prepare <case-id> --directory /tmp/<case-id>-workloadFor a standard native comparison, add --native-lane fresh to prepare and inspect
the resolved .github/actions/benchmark-phase/action.yml. Its cache publication
behavior and toolchain come from the case definition. Runtime canary inputs pass
through the shared workflow; they do not replace the case identity or recipe.
Fresh diagnostics without a series ID pass the cold action's cache_scope
output through the generated wrapper and job output into the warm action.
Both provider matrix jobs must publish the same cohort value. Keep this handoff
covered alongside declared-series execution; testing only a series can hide a
missing legacy input because the series scope is resolved independently.
Check prepared source, recipe, commands, and output verification. Recipe digest changes require review. Schema and workflow checks do not establish that a build works or that its comparison is fair.
bin/bench start <case-id> --series screening-01 --lane fresh --samples 2results/<case-id>/<series-id>/series.json freezes the question, providers, pins,
scope, phases, and sample count before execution. Use a new series for a changed
plan. Publication series normally use the case's declared ten measurements where
cost permits. Smaller screening series must be identified as screening.
Keep execution files at the reviewed harness ref when creating and dispatching the series, and commit the plan before execution. The definition digest includes shared execution and reporting code as well as the case. A later code change invalidates the plan for that checkout. Return to the original harness ref to continue its declared observations, or start a new series before execution. Adding result records does not change the execution definition.
Fresh series declare either a cold build and identical-source replay, or a cold
parent seed and changed-source build. The workflow's phases declares the latter
as ["cold", "commit"]; it must never be labelled an identical-source warm replay.
Rolling series contain changed-source commit observations and retain the parent cache seed.
Identities include case and series. Cases share a workspace but must not share
cold cache identity accidentally. A fresh tag or scope establishes a logical
cold seed; it does not establish an empty physical provider store. Keep existing
workspace/repository storage context and do not claim zero prior stored content
from namespace isolation alone.
Fresh samples run independently, including samples in the same series. Dispatch the declared sample numbers without waiting for earlier samples to finish. Each sample has a separate run/attempt scope. Its warm jobs still depend on its own cold jobs; they need the completed seed. Do not add a repository-wide queue.
Native rolling comparisons that update one cache series remain ordered. Their queue is scoped
to case and series, or the ref for legacy diagnostics. Deno and Zed retain their
dedicated Cargo-chain locks because their seed identity is shared across runs.
Docker-corpus rolling proofs remain diagnostic; their ordering and seed lineage
must be qualified before scheduling. They are not independent fresh samples.
Rolling groups use queue: max so newer requests do not replace waiting
observations. GitHub documents the 100-run queue limit.
scripts/lint-workflows.rb checks that property and runs Actionlint 1.7.12 for
the remaining workflow syntax; that linter version does not recognize queue.
GitHub runner limits can still cause queueing. Actions Cache quota remains repository-wide: record occupancy, eviction, and the actual restored key. Separate keys and concurrent execution do not establish independent quotas or seed availability. Keep fixture repositories when repository boundaries are part of the question. Queue and job duration remain context, outside the declared build-and-reuse comparison.
Use bin/bench run <case-id> --series <series-id> --sample 1 --ref <reviewed-ref> for a registered workflow.
Repeat with the next declared sample number. Use --workflow <filename> when a
case has more than one execution path. A receipt is retained for each sample and
workflow; a changed case requires a new series. Use --variant at series creation
when measuring a particular workload variant.
Workflows that declare variants require one of those variants at series creation.
When a variant selects a workflow input, declare variant_input in the workflow
entry. The plan fills that input from --variant and rejects a conflicting value.
The Zed layer proof uses the same combined cold seed for its target-only,
sccache-only, and combined restore probes; it establishes output correctness and
cache behavior rather than a comparison between providers. The frozen variant
selects cache_layer; one shared restore matrix runs only that selected probe.
Manual cache_layer: all runs the diagnostic matrix against a common cold seed.
Inspect the planned receipt before retrying an uncertain dispatch. Never
redispatch blindly after an interrupted request.
Time the operation declared by the case: build and cache reuse, with storage alongside it. Record restore/setup, build, and save when measured; arbitrary phase subdivisions are unnecessary. Unrelated dependency installation, runner queue, verification, and complete job duration remain context. When upstream combines them, use a reviewed boundary or label the combined measurement. Never call it compile time or subtract guessed dependency-install durations.
Canonical phase JSON includes case, series ID and sample, provider, lane/phase,
exact source, runner environment, verified output checks, primary measurement and
scope, storage with its measurement source, and evidence links. The canonical
Ruby reporter keeps legacy phase and lane fields for existing consumers. Series
records add explicit comparison fields; legacy records without them cannot
support a new series claim.
The series and dispatch retain the requested CLI release name. Product evidence
retains the version the CLI reports. A canary such as vcli-canary-<commit> can
report 1.34.0; those are different identifiers. Preserve the installer logs for
the selected artifact. The reported version alone does not identify a canary's
source commit. Released version requests must match the reported version, and
every BoringCache observation in a series must use the same reported CLI and Action.
In legacy phase JSON, a commit record carries lane: rolling, including the
changed-source step of a fresh Cargo proof. The declared series lane and phase
sequence identify that evaluation as fresh cold plus commit. The legacy
phase label does not establish a rolling chain or an identical-source replay.
The phase verification flag confirms output checks preceding that record.
verification.declared_checks lists the case's requirements; it does not claim
that checks for later phases have already run. Preserved job logs establish the
executed checks. Earlier imported records retain their original fields.
bin/bench record <case-id> --series screening-01 --record /tmp/phase.jsonRecording rejects duplicate slots. Reporting checks matched source, case, variant,
and runner environment across provider pairs. It reports every declared
observation, missing slots, medians, ranges, and measured storage counts. Missing
storage stays unmeasured. Reports require publication review.
Correctness proofs retain successful observations and execution checks without
a numeric correctness median or provider-comparison status. A storage comparison
uses the declared provider byte measurement for its primary median and range.
Reports retain all declared observations, including slow and failed runs.
Generated output contains measurements, sources, and check results. It does not
assign causes, classify a time difference as a tie, or recommend a provider.
Human or AI analysis is a separate activity; do not add it to generated reports
or the series catalog. Historical reviews remain at their original paths.
When review finds a measurement or fairness issue, keep the original plan and
records. Add methodology-review.json beside the series plan with
schema_version: 1, the exact plan_sha256, nonempty issues, and HTTPS
evidence_links. Reporting and the catalog retain all observations and mark the
series ineligible for a comparative claim. A corrected definition needs a new
series before execution; an old successful run must not acquire a revised scope.
Provider-reported bytes describe the selected cache, not total workspace usage
or billable storage. The BoringCache probe uses exact resolved tags and its
reported KV or archive size; the GitHub probe uses the selected cache key's
reported archive size. Keep those sources explicit. Do not infer equivalent
compression, cross-tag deduplication, or physical storage efficiency from these
counts alone.
Check coverage against the selected cache profile. For a combined Docker and
tool-cache profile, Docker-tag bytes alone do not establish total selected-profile
storage. Record the measured subset and missing coverage in the storage fields; do not
infer additional product tags in the harness.
GitHub storage reporting receives the read-only token only in the reporting step.
The selected exact-key rows are retained with their cache IDs, refs, versions,
timestamps, and byte fields. A missing row or byte field remains unmeasured;
an explicit zero-byte row is measured zero.
The selected storage total requires byte measurements for every resolved tag.
A partial probe retains its measured subtotal and unmeasured tags in
storage_breakdown; storage_bytes remains unmeasured. A missing tag or byte
field must not become a zero-byte measurement or a complete selected-cache total.
The breakdown retains the selected probe rows' status, identity, cache type, and
byte fields so coverage can be checked later. Archive entries use their archive
byte fields; an unrelated KV observation must not replace the archive size.
action.trust_state preserves the product's requested and resolved policy,
write permission, and decision source. Preserve the raw Action evidence as well;
do not infer publication permission from a successful build or a cache hit.
bin/bench preserve --repository boringcache/benchmarks --run <run-id> --directory /tmp/evidence-<run-id>
bin/bench verify-evidence --directory /tmp/evidence-<run-id>
bin/bench collect <case-id> --series screening-01 --sample 1 --directory /tmp/evidence-<run-id>collect checks the verified export against the retained dispatch, imports the
original phase artifacts, checks completion, and regenerates the report. It
rejects conflicting records before importing them and can resume an identical
import. It reads dedicated phase artifacts and canonical phase filenames inside
mixed proof bundles, retaining raw product JSON separately. It never derives
missing phase measurements from logs. record, finish,
and report remain available separately for reviewed historical imports.
Publish the verified bundle to durable storage with a stable link. Experimental
cache retention is not permanent evidence retention. The exporter includes all
available attempts and records gaps; partial exports remain partial. Keep
original execution URLs as provenance.
Cancelled requests remain in their original series. A run cancelled before any
jobs start can have a valid empty log ZIP and a complete scoped export, with no
measurements. Empty logs for a run that has jobs remain an evidence gap. Use a
new series after correcting an execution contract; do not replace the old request
or infer phase measurements from its logs.
finish checks the preserved run and job conclusions and post-step logs. A
green job with a reported BoringCache post-step failure remains unqualified.
Reports require a completion check for every run represented by their phase
records; complete timing records alone cannot qualify a comparison.
Retain the raw phase and product evidence when a later reporting step fails.
Summary artifact paths and names must include the reporter's selected variant
and lane. scripts/check-report-contract.rb checks that naming contract.
Review correctness, matched arms, storage semantics, sample completeness, and environment differences separately from the report. Publication decisions belong to that review. Website claims link to specific reports and preserved evidence.
The legacy data/latest/ collector remains a historical compatibility feed.
It reads original execution repositories and does not mix central Actions
summaries into those windows. Publish a central series through its canonical
report after evidence and methodology review; update the website's reviewed
evidence selection to an immutable report link. A green workflow or a case's
existing publication status does not approve a new series. Replacing the legacy
feed requires a separately qualified adapter for canonical series reports.
Canary monitoring follows canary_location in suites/published.json. Keep it
historical while the original repository owns the schedule. Set it to central
only when its central dispatcher and receipt have been qualified and its old
schedule retired. Missing, stale, or failed executions remain failures; importing
a definition does not move its active schedule.
Generate the common series catalog after recording or completing an evaluation:
bin/bench catalogdata/latest/series.json indexes every declared series through the same report
validator and calculations. It includes planned, requested, incomplete, failed,
completed, and invalid series, their sample counts, measurements, original runs,
completion errors, and scoped archive links. Requested means a dispatch receipt
exists; this offline catalog does not assert that GitHub queued or started a job.
Missing or stale stored reports are identified without rewriting them. Run
bin/bench report <case-id> --series <series-id> to generate the report after
reviewing new records. Catalog generation does not approve publication or merge
central observations into historical reporting windows.
Use bin/bench sync <case-id> --output /tmp/<case-id>-source-proposal.json in a
disposable harness checkout, or the manual source-sync.yml workflow. It inspects
the declared upstream branch, fetches the observed revision, checks fast-forward
history and recipe fidelity, and produces a proposal. A passing source proposal
does not verify the new build.
For verified source pairs, keep the candidate source environment from the proposal,
run its required build and output checks, then use bin/bench update-source <case-id> --directory <verified-candidate> to update the reviewed pins. Create a
new series whenever the frozen definition changes. Retain the prior series and
its evidence.
Changed-source comparisons must declare the exact seed and destination revisions and prove the intended cache continuity. Repeating one frozen source revision does not establish changed-source performance. Legacy rolling workflows remain diagnostic until their source sequence, seed lineage, and central caller have been qualified. Do not enable their schedule or promote their timing merely because the workflow completed. The series reporter keeps rolling observations but marks them invalid for comparison while seed lineage is unsupported.
Tool coverage distinguishes configured cases from missing
workload qualification. For a new adapter, declare its released .boringcache.toml
settings and pass its mode through the shared provider wrapper. Reuse preparation,
timing, scope, and reporting. Add a wrapper input only when the released Action
requires it for the case; do not copy every product option into a second API.
Verify the input against the Action metadata and product interface checks, then
run cold and reuse phases with output checks before adding a schedule.
The immutable boringcache/one pin in .github/actions/boringcache/action.yml is
the only maintained product Action pin. Its released default owns the ordinary
CLI version. Use the existing runtime cli_version input for an exact release or
canary experiment; do not add separate version constants to case actions.
For an Action or CLI update, read the released Action metadata and public plan contract, update the one pin when needed, then run the case, workflow, report, and product-owned interface checks. Start new screening series for affected adapters from a committed ref. Inspect output correctness, trust state, raw final evidence, native errors, storage coverage, and post-step completion before switching scheduled callers. A package download or passing schema check alone does not qualify the new product version.
Dependabot proposes weekly updates for the evidence hook's official GitHub SDKs. Guardrails install the exact lockfile, audit its dependencies, and execute the hook against the installed SDKs. SDK changes also require a live run that retains the original evidence after product cleanup. This dependency path does not update the BoringCache Action pin.
Keep previous series, pins, and observations unchanged. New fields in product evidence must survive in the raw artifact even before the canonical reporter uses them. Update the shared reporter when measurement semantics change, and retain explicit unknown or unmeasured values. Fix missing product capabilities in the product; do not reproduce its installers, cache identity, proxy lifecycle, or internal evidence normalization in individual benchmark cases.
Until managed native adapters are available, Moon, Pants, Buck2 and sbt use the
shared reapi-fresh-benchmark.yml workflow and reapi-benchmark action. This is
an explicitly selected direct CLI path. ci run owns renewable OIDC identity;
cache-registry --reapi-port owns cache serving, publication and shutdown flush.
The shared runner supervises that command and retains its exit status and logs.
Client settings live in one shared helper; cases contain workload recipes and
reviewed configuration patches. Do not add per-case proxy wrappers.
This path downloads the selected release asset and verifies its published SHA-256 through one shared helper. It retains the requested release, asset and checksum. It records the actual CLI version without inventing One Action refs or an Action evidence envelope. A managed product adapter should replace this shared entrypoint when available.
The bazel-remote comparison arm transfers its store through a GitHub artifact outside timing. These series declare correctness as the primary metric and must not support an end-to-end provider speed comparison. Native hit evidence and matching outputs on a separate warm worker are required. Missing storage stays unmeasured. New performance comparisons require a separately reviewed service and transfer measurement boundary.