Summary
Compiled w8a16 model is reported as w16a16
Priority: P1, confirmed by the release driver on 2026-09-23.
Bugbash reference: 2026-09-22 / BUG-004.
Status: Observed in the first test round; no fix or post-fix regression is recorded. This issue is based on retained test evidence, not a new test run.
Environment
- Test date: 2026-09-22; findings triaged on 2026-09-23.
- CLI: winml-cli 0.4.0, installed from a local wheel. The exact source commit/build provenance is not recorded in the test evidence.
- Hardware: Snapdragon X Elite X1E80100, Qualcomm Hexagon NPU.
- OS: Windows 11 26H1, build 28000.2956, ARM64.
- Python: 3.11.15, workspace virtual environment.
- EP: QNNExecutionProvider from WinML Catalog, package 2.2480.49.0.
- ONNX Runtime: 1.27.1.202607110137, Windows ML distribution.
- ONNX: 1.18.0; PyTorch: 2.14.0; Transformers: 4.57.6.
- Model: catalog microsoft/resnet-50, image-classification, pixel_values FP32 [1, 3, 224, 224].
- Quantization, where applicable: w8a16, uint8 weights / uint16 activations, 10 calibration samples.
Reproduction
Run in PowerShell with winml-cli 0.4.0 installed. Use a new output directory and retain the external files generated alongside the compiled ONNX.
$results = '.\repro-compiled-precision'
New-Item -ItemType Directory -Path $results | Out-Null
uv run winml export -m microsoft/resnet-50 -o "$results\resnet.onnx"
uv run winml optimize -m "$results\resnet.onnx" --ep qnn --device npu --disable-ort-graph-optimization -o "$results\resnet-opt-no-ort.onnx"
uv run winml quantize -m "$results\resnet-opt-no-ort.onnx" -p w8a16 --samples 10 --task image-classification --model-id microsoft/resnet-50 -o "$results\resnet-w8a16.onnx"
uv run winml compile -m "$results\resnet-w8a16.onnx" --ep qnn --device npu -o "$results\resnet-compiled.onnx"
uv run winml perf -m "$results\resnet-w8a16.onnx" --ep qnn --device npu --iterations 10 --warmup 2 -o "$results\precision-raw.json"
uv run winml perf -m "$results\resnet-compiled.onnx" --ep qnn --device npu --iterations 20 --warmup 3 -o "$results\precision-compiled.json"
Compare model_info.precision in the two reports. The optimize flag in preparation bypasses a separate serialization issue.
Expected result
Preserve the source precision, or report it as unknown when compiled weights are opaque.
Actual result
Both perf commands exit successfully, but their model_info.precision values disagree:
| Input |
Reported precision |
| Raw uint8-weight / uint16-activation QDQ model |
w8a16 |
| Compiled EPContext from that same model |
w16a16 |
This discrepancy was independently confirmed from the saved JSON reports.
Workaround
Use the verified pre-compile quantization metadata when interpreting results; there is no verified CLI workaround that corrects the compiled report field.
Scope and evidence
The compiled outer ONNX exposes UINT16 activation-side Q/DQ initializers, while weights are opaque in the external EPContext. This confirms a report-label defect, not an actual conversion of weights to 16 bits. Successful inference/exit code does not make the precision field correct. Preserve known source precision or report unknown when weight precision cannot be established.
The following evidence is retained locally by the reporter; these filenames are an index, not uploaded attachments:
quantize-w8a16.stdout.log
perf-raw.json
perf-compiled.json
onnx-validation.json
Summary
Compiled w8a16 model is reported as w16a16
Priority: P1, confirmed by the release driver on 2026-09-23.
Bugbash reference: 2026-09-22 / BUG-004.
Status: Observed in the first test round; no fix or post-fix regression is recorded. This issue is based on retained test evidence, not a new test run.
Environment
Reproduction
Run in PowerShell with winml-cli 0.4.0 installed. Use a new output directory and retain the external files generated alongside the compiled ONNX.
Compare model_info.precision in the two reports. The optimize flag in preparation bypasses a separate serialization issue.
Expected result
Preserve the source precision, or report it as unknown when compiled weights are opaque.
Actual result
Both perf commands exit successfully, but their model_info.precision values disagree:
This discrepancy was independently confirmed from the saved JSON reports.
Workaround
Use the verified pre-compile quantization metadata when interpreting results; there is no verified CLI workaround that corrects the compiled report field.
Scope and evidence
The compiled outer ONNX exposes UINT16 activation-side Q/DQ initializers, while weights are opaque in the external EPContext. This confirms a report-label defect, not an actual conversion of weights to 16 bits. Successful inference/exit code does not make the precision field correct. Preserve known source precision or report unknown when weight precision cannot be established.
The following evidence is retained locally by the reporter; these filenames are an index, not uploaded attachments:
quantize-w8a16.stdout.logperf-raw.jsonperf-compiled.jsononnx-validation.json