Skip to content

feat: gate quantization and speculative decoding on measured evidence - #32

Merged
github-actions[bot] merged 1 commit into
mainfrom
feat/19-engine-variants
Oct 3, 2026
Merged

github-actions[bot] merged 1 commit into
mainfrom
feat/19-engine-variants

Conversation

@Yash-Chindam

Copy link
Copy Markdown
Owner

Fifth PR closing gaps between the design spec and the implementation.

Gap this closes

§9: quantization formats "as independent variants" measured for memory, throughput, latency and quality; speculative decoding "as an experiment"; and the design target "documented quality delta for every optimization variant". Quantization was one field on a model card and speculative decoding appeared nowhere.

What changed

  • EngineVariant in the catalog (variants:), bound to one immutable base revision, of kind quantization or speculative-decoding.
  • In development, a variant renders under experiments in config/ray-serve.yaml as its own application with no warm replica.
  • A variant cannot reach staging/production without a baseline and a variant benchmark. Runs that differ in dataset, workload, hardware, or concurrency are rejected at load.
  • variant_verdict: quality loss beyond tolerance is always a regression; speculative decoding must lower p95 or raise throughput; quantization must reduce measured GPU memory.
  • A promoted variant is applied to the base model's engine arguments and added to the cache fingerprint.
  • GET /v1/registry/variants lists every variant with its deltas and regressions.
  • python -m llm_router.evaluation runs a dataset against an OpenAI-compatible endpoint and prints a catalog benchmark record, reading GPU memory and draft acceptance from the engine's /metrics when present.

Not done

  • Nothing has been measured. Both committed variants (general-gptq, high-capability-speculative) are in development with no evidence. Producing it needs a GPU and a running vLLM engine.
  • The benchmark CLI wrapper (main) is excluded from coverage and has not been run against a real engine; the transport, measurement parsing, and record shaping underneath it are tested with a mocked endpoint.
  • The CLI sends requests one at a time, so its throughput figure is sequential.

Test plan

  • 21 new tests: promotion blocked on quality loss, on speculative overhead, on unmeasured or unreduced memory, and on non-comparable runs; experiment rendering; cache identity on promotion; committed ray-serve.yaml matches the catalog
  • ruff format --check ., ruff check ., mypy clean
  • pytest tests/unit tests/integration: 262 passed, coverage 98%
  • Playwright end-to-end suite: 9 passed locally

🤖 Generated with Claude Code

Section 9 treats quantization formats as independent variants to be
measured for memory, throughput, latency and quality, and speculative
decoding as an experiment, because low draft-token acceptance can add
overhead. The design targets require a documented quality delta for every
optimization variant. Quantization was a field on a model card and
speculative decoding did not exist.

An engine variant is now declared in the catalog against one immutable
base revision. In development it is rendered as its own never-warm
application beside the base model so it can be benchmarked without taking
traffic. It cannot be staged or promoted without a baseline and a variant
benchmark, and the catalog refuses runs that differ in dataset, workload,
hardware or concurrency.

A quality loss beyond tolerance blocks promotion whatever the gain.
Speculative decoding must also be faster, and a quantized variant must
show the memory it saves. A promoted variant is written into the base
model's engine arguments and changes the cache identity.

python -m llm_router.evaluation benchmarks a live endpoint and prints the
catalog record. Neither committed variant has been measured, so both stay
in development.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation area/api area/tests labels Oct 3, 2026
@github-actions
github-actions Bot merged commit 2e58ae8 into main Oct 3, 2026
6 checks passed
@github-actions
github-actions Bot deleted the feat/19-engine-variants branch October 3, 2026 14:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/api area/tests documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant