Skip to content

feat(xtest): cross-SDK benchmark roll-up and decision-first summaries - #623

Draft
dmihalcik-virtru wants to merge 1 commit into
DSPX-4372-03-payload-sizesfrom
DSPX-4372-04-reporting-rollup
Draft

dmihalcik-virtru wants to merge 1 commit into
DSPX-4372-03-payload-sizesfrom
DSPX-4372-04-reporting-rollup

Conversation

@dmihalcik-virtru

Copy link
Copy Markdown
Member

Split out of #583. Stacked on #622.

Each SDK's bench job already reports its own verdict; nothing combined the three into one answer to "did this change regress anything," and a reviewer had to open three job summaries to find out.

  • perf/report.py: the per-SDK summary is reordered for progressive disclosure (TL;DR, compared builds, what changed, then the full table), and the effect-at-a-glance scale is tail-compressed so a planted outlier does not flatten every ordinary-sized effect. Renders bake-off results from the K-arm core alongside gated verdicts.
  • perf/aggregate.py (new): pure-stdlib cross-SDK roll-up read from the JSON artifacts, not the markdown, so a partially failed matrix still gets a combined answer for whichever SDKs reported.
  • xtest.yml: new benchmark-summary job downloads every SDK's bench-result-* artifact and publishes the roll-up as the step summary once all three matrix jobs finish. Renamed the per-SDK upload from the emoji-prefixed bench-<sdk> to bench-result-<sdk>, which is what the roll-up's artifact-name pattern actually matches.

Built on #622 (payload sizes) and #621 (K-arm core, for bake-off rendering).

Each SDK's bench job already reports its own verdict; nothing combined the
three into one answer to 'did this change regress anything', and a reviewer
had to open three job summaries to find out.

- perf/report.py: the per-SDK summary is reordered for progressive
  disclosure -- TL;DR, compared builds, what changed, then the full table --
  and the effect-at-a-glance scale is tail-compressed so a planted 3x outlier
  does not flatten every ordinary-sized effect into the same pixel. Renders
  bake-off results (from the K-arm core) alongside the gated verdicts.
- perf/aggregate.py (new): pure-stdlib -- no dependency sync needed in a job
  that runs after everything else already finished -- cross-SDK roll-up
  read from the JSON artifacts, not the markdown, so a partially failed
  matrix still gets a combined answer for whichever SDKs reported.
- xtest.yml: new benchmark-summary job downloads every SDK's bench-result-*
  artifact and publishes the roll-up as the step summary once all three
  matrix jobs finish. Renamed the per-SDK upload from the emoji-prefixed
  'bench-<sdk>' to 'bench-result-<sdk>', which is what the roll-up's
  artifact-name pattern actually matches -- the emoji form was never
  discoverable by it.

Built on the K-arm core (bake-off rendering) and payload sizes (the summary
reads whatever payload set the run measured).
@coderabbitai

coderabbitai Bot commented Sep 23, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sonarqubecloud

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant