Skip to content

[executorch][cuda] Time AOTI autotune candidates with CUDA graphs so a busy CPU does not decide the pick - #23510

Open
Gasoonjia wants to merge 2 commits into
gh/gasoonjia/192/basefrom
gh/gasoonjia/192/head
Open

Gasoonjia wants to merge 2 commits into
gh/gasoonjia/192/basefrom
gh/gasoonjia/192/head

Conversation

@Gasoonjia

@Gasoonjia Gasoonjia commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

AOTInductor picks, at compile time, the config of every Triton kernel it generates, the implementation of every matmul (its Triton templates), and the config of every user triton.autotune kernel (our custom ops, e.g. triton::sdpa*). It does so by timing the candidates. This diff makes that timing reliable while the CPU is saturated. The code is in a new package, backends/cuda/autotune/ (cuda_graph_timing.py).

  • How candidates are timed today. Every candidate timing goes through torch._inductor.runtime.benchmarking.benchmarker, which dispatches on device type through a registry. The default CUDA timing brackets each call with a pair of host-recorded events.
  • Why that fails during an export. The CPU is saturated, because Inductor compiles kernels in parallel worker processes. Under that load the readings become bimodal: a few-microsecond kernel reads either its own time or tens of microseconds more. The autotuner then regularly picks configs several times slower than the best.
  • What this diff does. For the duration of the CUDA backend's AOTI compile, it registers a CUDA timing that captures several calls of the candidate into one CUDA graph. Each call is preceded by an L2 flush, as Inductor's default timing does, and bracketed by events recorded inside the graph. A replay is a single submission, so a busy CPU cannot add to the measured time.
  • Keeping captures succeeding:
    • each graph and its memory pool are released right after timing;
    • on CUDA OOM the allocator cache is emptied and the capture retried;
    • if that fails too, the model weights moved to the GPU for the compile are parked on the CPU while the candidate is timed. AOTI autotunes on fabricated inputs, so the weights are not needed for this;
    • a candidate that cannot be captured falls back to Inductor's default timing.
  • Default and opt-out. On by default; the new cuda_graph_autotune_timing compile spec (ON/OFF) turns it off.

Results (A100), CPU-saturated toy export (each pick scored by re-timing every candidate with the CPU paused):

default this diff
mean / worst regret of the picks 1.07x to 1.95x / up to 5.5x 1.001x / 1.014x

The diff above (representative inputs for data-dependent kernel arguments) has the solo muse-glimmer A/B of both diffs together.

Differential Revision: D123563417

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23510

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures, 1 Cancelled Job

As of commit feb2a4c with merge base 8789aa5 (image):

NEW FAILURES - The following jobs have failed:

CANCELLED JOB - The following job was cancelled. Please retry:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

[ghstack-poisoned]
@Gasoonjia
Gasoonjia deployed to upload-benchmark-results October 7, 2026 00:17 — with GitHub Actions Active

This branch was successfully deployed

2 active deployments
upload-benchmark-results — feb2a4cb Deployed Oct 7, 2026 by Gasoonjia via upload-benchmark-results #20418
cadence — feb2a4cb Deployed Oct 6, 2026 by Gasoonjia via hifi-op-test / hifi4 #32077
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant