Repository navigation
feat: gate quantization and speculative decoding on measured evidence - #32
Merged
Merged
Conversation
Section 9 treats quantization formats as independent variants to be measured for memory, throughput, latency and quality, and speculative decoding as an experiment, because low draft-token acceptance can add overhead. The design targets require a documented quality delta for every optimization variant. Quantization was a field on a model card and speculative decoding did not exist. An engine variant is now declared in the catalog against one immutable base revision. In development it is rendered as its own never-warm application beside the base model so it can be benchmarked without taking traffic. It cannot be staged or promoted without a baseline and a variant benchmark, and the catalog refuses runs that differ in dataset, workload, hardware or concurrency. A quality loss beyond tolerance blocks promotion whatever the gain. Speculative decoding must also be faster, and a quantized variant must show the memory it saves. A promoted variant is written into the base model's engine arguments and changes the cache identity. python -m llm_router.evaluation benchmarks a live endpoint and prints the catalog record. Neither committed variant has been measured, so both stay in development. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fifth PR closing gaps between the design spec and the implementation.
Gap this closes
§9: quantization formats "as independent variants" measured for memory, throughput, latency and quality; speculative decoding "as an experiment"; and the design target "documented quality delta for every optimization variant". Quantization was one field on a model card and speculative decoding appeared nowhere.
What changed
EngineVariantin the catalog (variants:), bound to one immutable base revision, of kindquantizationorspeculative-decoding.development, a variant renders underexperimentsinconfig/ray-serve.yamlas its own application with no warm replica.staging/productionwithout a baseline and a variant benchmark. Runs that differ in dataset, workload, hardware, or concurrency are rejected at load.variant_verdict: quality loss beyond tolerance is always a regression; speculative decoding must lower p95 or raise throughput; quantization must reduce measured GPU memory.GET /v1/registry/variantslists every variant with its deltas and regressions.python -m llm_router.evaluationruns a dataset against an OpenAI-compatible endpoint and prints a catalog benchmark record, reading GPU memory and draft acceptance from the engine's/metricswhen present.Not done
general-gptq,high-capability-speculative) are indevelopmentwith no evidence. Producing it needs a GPU and a running vLLM engine.main) is excluded from coverage and has not been run against a real engine; the transport, measurement parsing, and record shaping underneath it are tested with a mocked endpoint.Test plan
ray-serve.yamlmatches the catalogruff format --check .,ruff check .,mypycleanpytest tests/unit tests/integration: 262 passed, coverage 98%🤖 Generated with Claude Code