Repository navigation
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23514
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New Failure, 2 Cancelled Jobs, 1 Unrelated FailureAs of commit 6907db4 with merge base 8789aa5 ( NEW FAILURE - The following job has failed:
CANCELLED JOBS - The following jobs were cancelled. Please retry:
UNSTABLE - The following job is marked as unstable, possibly due to flakiness on trunk:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Oct 6, 2026
This PR needs a
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack from ghstack (oldest at bottom):
Decode-sized (M <= 4) INT6 linears (
CudaDp4aPlanarInt6Tensor, GGUF Q6_K) move ontoQuantizedGemmFamilyastriton::int6_quantized_gemm_m{1,2,3,4}, the same way as INT4 in the parent diff. This removes theint6_plain_mmC shim everywhere.Kernels,
triton/kernels/int6_quantized_gemm.py. W6A8 DP4A.quantize_activations_q8.ql/qh, the constant -32 offset is folded into the INT32 dot product through a DP4A activation sum, and both scale levels are applied in FP32.quantized_gemm_utils.py/quantized_gemm_family.py: activation quantization, DP4A/warp-sum helpers, split-K rule and reduce, the autotune space and pruning, the legality checks, and the launch skeleton. Only the INT6 decode, the main kernels and the INT6 rules live here.Dispatch,
quantize_op_dispatch/int6_dispatch.py. It uses the sharedquantized_linearandchunked_dequant_linear(the INT6 dequant is now chunked along N like the others). Unsupported inputs fall back to dequant +F.linearand never raise. Theint6_plain_mmschema and its Meta/CUDA impls are removed.C shim removal (fbcode + xplat):
runtime/shims/int6_plain_mm.{h,cu,cuh}, its gtest and its benchmark;runtime/targets.bzl,CMakeLists.txtand the shim tests'CMakeLists.txt;gen_plain_mm_test_vectors.py;custom_ops_to_c_shimsentries incuda_backend.py, and thetest_sort_shimexpectations;triton.int6_quantized_gemm_m1in the graph.Op level (A100; full sequence = activation quantization + GEMM (+ reduce); the shim built from the pre-diff sources; each case runs the best config of the pruned space, then shim and Triton alternate for 7 rounds and medians are compared). All five Q6_K linear shapes of the gemma4_31b, muse-glimmer and dflash GGUFs, M = 1..4, plus the 4-row bucket at a dynamic M = 2, 3:
Geomean 1.134x for static M and 1.174x for dynamic M. The one case below 1.0x is gemma attn_v at M = 3: 0.993x here, and a 0.996x tie in KernelAgent's own measurement. Mean relative difference vs the shim is < 0.0005 everywhere.
SEE_AB_TABLE
Differential Revision: D123644066