Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Quantized LLM Serving: BF16 vs FP8 on H100

An end-to-end study that connects quantization error to production serving behavior. It implements symmetric quantization from first principles, serves the same 7B model with vLLM in BF16 and FP8, and benchmarks both configurations with GuideLLM.

Headline results

Measured on one NVIDIA H100 80 GB with Qwen/Qwen2.5-7B-Instruct:

Metric BF16 FP8 Change
Logged model weights 14.29 GiB 8.17 GiB -42.8%
Saturated output throughput 4,804.8 tok/s 7,159.2 tok/s +49.0%
Synchronous median request latency 1.581 s 1.096 s -30.6%
Synchronous median inter-token latency 6.11 ms 4.20 ms -31.3%
Synchronous median TTFT 21.59 ms 24.51 ms +13.6%

BF16 versus FP8 serving results

FP8 helped decode more than prefill. That is consistent with token-at-a-time decode repeatedly streaming weights and KV-cache data, while prompt prefill has higher arithmetic intensity. The 1.49x throughput gain is meaningful, but below the idealized 2x implied by halving weight precision because weights are not the only serving cost.

What is in the project

  • notebooks/quantization_and_serving.ipynb — quantize/dequantize implementation, error analysis, vLLM launch configuration, benchmark runner, and interpretation.
  • results/q1/ — signal-to-noise and error measurements from 3- to 8-bit symmetric quantization.
  • results/q2/ — the compact comparison plus the four machine-readable GuideLLM reports used to calculate the table above.
  • tests/ — dependency-free checks for artifact consistency and headline metrics.

Quantization experiment

Per-channel scaling was substantially more accurate than one scale for the whole tensor. At 8 bits, measured SNR improved from 18.09 dB per tensor to 42.65 dB per channel. At 4 bits, the same comparison was 5.15 dB versus 17.71 dB.

SNR by bit width

Reproduce

The analysis cells run on CPU. Reproducing the serving experiment requires an NVIDIA GPU supported by vLLM; the recorded FP8 results require Hopper-class hardware.

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
jupyter lab notebooks/quantization_and_serving.ipynb

Run the lightweight artifact checks without installing GPU dependencies:

python -m unittest discover -s tests -v

Limitations

  • The serving runs are short, synthetic 30-second benchmarks rather than a trace replay from production.
  • GuideLLM measures serving behavior, not answer quality. FP8 should also be evaluated on a representative golden set before deployment.
  • The synchronous samples are small, so tail-latency estimates are directional.
  • Results are specific to this model, vLLM/GuideLLM versions, and H100 system.

Provenance

This project was developed from a performance-engineering course exercise. The implementation, measured runs, analysis, and portfolio packaging are the author's work; the notebook retains some supplied experimental harness code. See NOTICE.md.

About

Measured BF16 vs FP8 serving study on H100: quantization error, vLLM memory, latency, and throughput.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages