docs: framing matrix complete — a distribution property, not quantization - #214
Merged
Merged
Conversation
…tion The single-vs-batched test run at every endpoint: Python ORT 480/480 at fp32, dynamic int8, AND static int8; onnxruntime-node 89-101/485 at every precision including fp32. Static scales do not repair the node build - its kernels differ across batch shapes before quantization enters. Shipped paths are single-framing so published byte parity is unaffected; framing is the parity variable of mixed-shape serving. Paper C 5.1 carries the final scoping; static int8 stays a candidate for the export path on speed and quality, decided by full-set DER.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The framing test run at every endpoint (same machine, ORT 1.23, 5 golden rows x 96 steps):
The instability is a property of the onnxruntime-node kernel paths (batch-shape-dependent numerics at fp32, before quantization enters), absent from the Python build at every precision. Static activation scales - the pre-registered fix - do not repair it; TODO.impl/11 closes with that negative plus a positive byproduct (static int8 is quality-clean on sample and ~8% faster than dynamic on CPU; candidate for the export path on its own merits, gated by full-set DER).
Paper C 5.1 carries the final scoping. Shipped paths are single-framing by construction, so published byte parity is unaffected; framing is the parity variable of mixed-shape serving (why the speculative tier degraded and was pulled).