Quantization from first principles + real-world vLLM serving benchmarks — BF16 vs on-the-fly FP8/INT8, measured with guidellm for memory, throughput, and latency trade-offs.
-
Updated
Jul 31, 2026 - Jupyter Notebook
Quantization from first principles + real-world vLLM serving benchmarks — BF16 vs on-the-fly FP8/INT8, measured with guidellm for memory, throughput, and latency trade-offs.
Artifact-backed LLM serving performance lab for vLLM baselines, official metrics, GuideLLM checks, and SGLang/PD scaffolding
Results parser for guideLLM with OpenSearch indexing capabilities
Add a description, image, and links to the guidellm topic page so that developers can more easily learn about it.
To associate your repository with the guidellm topic, visit your repo's landing page and select "manage topics."