An end-to-end study that connects quantization error to production serving behavior. It implements symmetric quantization from first principles, serves the same 7B model with vLLM in BF16 and FP8, and benchmarks both configurations with GuideLLM.
Measured on one NVIDIA H100 80 GB with Qwen/Qwen2.5-7B-Instruct:
| Metric | BF16 | FP8 | Change |
|---|---|---|---|
| Logged model weights | 14.29 GiB | 8.17 GiB | -42.8% |
| Saturated output throughput | 4,804.8 tok/s | 7,159.2 tok/s | +49.0% |
| Synchronous median request latency | 1.581 s | 1.096 s | -30.6% |
| Synchronous median inter-token latency | 6.11 ms | 4.20 ms | -31.3% |
| Synchronous median TTFT | 21.59 ms | 24.51 ms | +13.6% |
FP8 helped decode more than prefill. That is consistent with token-at-a-time decode repeatedly streaming weights and KV-cache data, while prompt prefill has higher arithmetic intensity. The 1.49x throughput gain is meaningful, but below the idealized 2x implied by halving weight precision because weights are not the only serving cost.
notebooks/quantization_and_serving.ipynb— quantize/dequantize implementation, error analysis, vLLM launch configuration, benchmark runner, and interpretation.results/q1/— signal-to-noise and error measurements from 3- to 8-bit symmetric quantization.results/q2/— the compact comparison plus the four machine-readable GuideLLM reports used to calculate the table above.tests/— dependency-free checks for artifact consistency and headline metrics.
Per-channel scaling was substantially more accurate than one scale for the whole tensor. At 8 bits, measured SNR improved from 18.09 dB per tensor to 42.65 dB per channel. At 4 bits, the same comparison was 5.15 dB versus 17.71 dB.
The analysis cells run on CPU. Reproducing the serving experiment requires an NVIDIA GPU supported by vLLM; the recorded FP8 results require Hopper-class hardware.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
jupyter lab notebooks/quantization_and_serving.ipynbRun the lightweight artifact checks without installing GPU dependencies:
python -m unittest discover -s tests -v- The serving runs are short, synthetic 30-second benchmarks rather than a trace replay from production.
- GuideLLM measures serving behavior, not answer quality. FP8 should also be evaluated on a representative golden set before deployment.
- The synchronous samples are small, so tail-latency estimates are directional.
- Results are specific to this model, vLLM/GuideLLM versions, and H100 system.
This project was developed from a performance-engineering course exercise. The implementation, measured runs, analysis, and portfolio packaging are the author's work; the notebook retains some supplied experimental harness code. See NOTICE.md.

