Benchmark-driven configuration optimization for vLLM inference servers
-
Updated
Sep 23, 2026 - Python
Benchmark-driven configuration optimization for vLLM inference servers
Auto-tuning for vllm. Getting the best performance out of your LLM deployment (vllm+guidellm+optuna)
Artifact-backed LLM serving performance lab for vLLM baselines, official metrics, GuideLLM checks, and SGLang/PD scaffolding
Quantization from first principles + real-world vLLM serving benchmarks — BF16 vs on-the-fly FP8/INT8, measured with guidellm for memory, throughput, and latency trade-offs.
Reproducible LLM serving benchmarks on Akamai Cloud GPUs. Terraform, cloud-init, and guidellm compare vLLM, SGLang, TGI, Ollama, llama.cpp, and NVIDIA NIM on latency, throughput, and cost per million tokens.
Measured BF16 vs FP8 serving study on H100: quantization error, vLLM memory, latency, and throughput.
Results parser for guideLLM with OpenSearch indexing capabilities
To associate your repository with the guidellm topic, visit your repo's landing page and select "manage topics."