LLM Inference Optimization: Quantization, KV Cache & Batching

Quick answer
To optimize a production LLM, first measure a fixed-model baseline. Then reduce memory pressure with an appropriate weight or KV-cache format, improve scheduling with continuous batching, test speculative decoding only on workloads where draft-token acceptance is healthy, and choose local or cloud deployment according to privacy, cost, reliability, and SLO requirements.
The useful sequence is measure → change one variable → measure again → check quality → canary → roll back if necessary. A smaller model or lower-precision model is not automatically faster if its runtime lacks an efficient kernel or if it causes retries and quality failures.
The SPEED framework
This framework complements the ROUTE framework for model routing. Routing decides which model should answer; inference optimization decides how that model is served.
S — Quantization that respects quality
Quantization represents weights, activations, or cached keys and values with fewer bits or a lower-precision floating-point format. It can reduce VRAM use and sometimes improve throughput, but the result is workload-dependent. A useful deployment decision must include both system metrics and task quality.
What should be quantized?
| Target | Potential benefit | Main risk | Validation |
|---|---|---|---|
| Weights | Lower model memory and often lower bandwidth pressure | Quality loss, especially on sensitive layers or tasks | Task accuracy, structured-output validity, VRAM |
| Activations | Lower memory traffic in supported kernels | Runtime and calibration incompatibility | Prefill latency and end-to-end quality |
| KV cache | More context tokens per GPU | Attention-quality loss or backend limitations | Long-context quality, OOM rate, TPOT |
INT4 AWQ or GPTQ can be valuable for some decoder-only models, while INT8 or FP8 may provide a safer quality/performance trade-off. Do not describe INT8 as universally safe or INT4 as universally faster. The implementation, group size, calibration data, and hardware kernel matter.
For current deployment details, consult the Hugging Face quantization documentation and TensorRT-LLM documentation. Record the exact format and runtime version in every benchmark.
P — Page and persist the KV cache

During autoregressive decoding, the model reuses keys and values from previous tokens. A paged cache stores those values in fixed-size blocks and maps logical sequences to physical memory. This reduces fragmentation and makes multi-request serving easier to schedule. The original PagedAttention paper describes the foundation, while the current vLLM documentation notes that its historical kernel explanation does not describe every detail of the current implementation.
Prefix reuse needs privacy boundaries
Prefix caching can reuse a shared system prompt or a stable document prefix, but the cache key must include the model version, tokenizer or template version, tenant, permissions, and relevant policy state. Never share a cached prefix across users merely because the text looks similar. Add TTLs, invalidation rules, and access checks.
KV quantization and offload
vLLM documents FP8 KV-cache options with calibration and layer-skipping controls. The benefit is more tokens in memory, not a guaranteed reduction in every latency metric. CPU offload can prevent an out-of-memory failure, but PCIe transfer and cache thrashing can make a request slower. Disk-backed cache should be treated as an exceptional capacity mechanism, not a default decode optimization.
E — Continuous batching and latency discipline
Continuous or in-flight batching admits and retires requests during the decode loop. It can improve aggregate utilization and throughput when requests have different lengths, but it can also increase queueing or tail latency if token budgets and scheduling priorities are poorly chosen.
| Metric | Meaning | Why it matters |
|---|---|---|
| TTFT | Time to first token | Prompt processing and queueing experience |
| TPOT / ITL | Time per output token or inter-token latency | Streaming smoothness after the first token |
| Throughput | Output tokens per second across requests | Capacity and unit economics |
| p95 / p99 | Tail latency percentiles | Reliability for the slowest users |
| Goodput | Throughput under explicit SLO limits | Prevents fast-but-unusable benchmarks |
Set a maximum number of new tokens, separate interactive and batch queues where appropriate, and report concurrency. A single tokens-per-second number without prompt length, output length, and concurrency is not a meaningful production comparison.
E — Speculative decoding: test acceptance, not hype
Speculative decoding uses a smaller proposer to draft several tokens and a larger target model to verify them. The original research describes a lossless sampling method under its assumptions, but the speedup depends on draft cost, token acceptance, target-model bottlenecks, sampling settings, and the serving workload.
The current vLLM guidance focuses on reducing inter-token latency in medium-to-low-QPS, memory-bound workloads and lists different proposer methods. Therefore, a fixed recommendation such as “always pair a 2–4B draft with a 7–13B target” is too broad. Test the pair on your own prompts.
Reproducible experiments you can run
The following experiments are designed to produce real measurements on your hardware. They intentionally do not claim a universal result. Publish the output with the GPU, driver, CUDA version, model revision, runtime version, prompt set, concurrency, and sampling configuration.
Experiment 1: establish a baseline
# Example vLLM server; pin the versions in your own environment
vllm serve MODEL_ID \
--dtype auto \
--max-model-len 8192 \
--gpu-memory-utilization 0.90
# Run the official benchmark tool available in your installed vLLM version
python benchmarks/benchmark_serving.py \
--backend vllm \
--model MODEL_ID \
--dataset-name random \
--random-input-len 1024 \
--random-output-len 256 \
--num-prompts 200 \
--request-rate inf
Record TTFT, TPOT, output tokens per second, p50, p95, p99, GPU memory, and request failures. Repeat at concurrency 1, 4, 16, and the expected production level. Store the command and raw output beside the published table.
Experiment 2: compare precision without hiding quality loss
# Run the same prompt set against baseline and quantized variants
python evaluate_prompts.py \
--model baseline \
--prompts prompts.jsonl \
--seed 42 \
--output baseline.jsonl
python evaluate_prompts.py \
--model quantized \
--prompts prompts.jsonl \
--seed 42 \
--output quantized.jsonl
Your prompt set should include extraction, reasoning, RAG, tool-use, refusal, and long-context cases. Compare exact match, JSON validity, citation coverage, task-specific score, and human review on failure slices. A lower perplexity score alone is not enough for a product decision.
Experiment 3: test speculative decoding
Keep the target model, prompt set, output limit, and sampling settings fixed. Change only the proposer configuration. Report proposer acceptance rate, TPOT, end-to-end latency, throughput, GPU memory, and quality equivalence. Test low-QPS and high-QPS separately because a latency win at low concurrency may not be a throughput win under a full queue.
D — Deployment roadmap
- Baseline first. Define SLOs for TTFT, TPOT, p95 latency, throughput, quality, cost, and error rate before changing the stack.
- Reduce memory pressure. Test weight precision and KV-cache settings separately. Keep a rollback configuration and a quality canary set.
- Improve scheduling. Add continuous batching, token budgets, queue priorities, and observability. Watch p95 and queue time, not only GPU utilization.
- Test speculative decoding. Measure acceptance and target SLOs at the traffic levels that matter. Remove it if the total system becomes slower or more expensive.
- Choose deployment deliberately. Compare local GPUs, cloud GPUs, and managed inference using total cost of ownership, data residency, reliability, and operational effort.
Inference optimization also connects to AI observability, agent memory design, and zero-trust access control. A faster server that leaks data or cannot explain a regression is not production-ready.
Common failure modes
| Failure | Why it happens | Control |
|---|---|---|
| INT4 looks fast but quality drops | Calibration or kernel does not fit the task | Slice-level quality canary and rollback |
| Batching raises p99 | Long requests starve interactive traffic | Token caps, queue separation, priority limits |
| Speculation adds overhead | Low draft acceptance or high QPS | Measure acceptance and disable by policy |
| Prefix cache leaks context | Cache key ignores tenant or permissions | Isolation, ACL checks, TTL, invalidation |
| Benchmark cannot be reproduced | Missing versions and workload details | Publish config, prompts, raw output, and seed |
FAQ
What is the fastest optimization to try?
Start with measurement and a memory profile. If the model fits comfortably and the workload is low-concurrency, quantization may be more useful than speculative decoding. If the GPU is underfilled by many concurrent requests, scheduling may have the larger effect.
Does quantization always reduce cost?
No. It can reduce memory and allow more requests per GPU, but a slower kernel, quality failures, retries, or a more expensive validation path can cancel the saving.
Should I use CPU or disk offload?
Use offload as a capacity and resilience option after measuring transfer cost. It is not a free latency optimization.
How do I make the article's benchmark credible?
Publish the exact model revision, runtime and driver versions, hardware, prompt distribution, input and output lengths, concurrency, sampling settings, warm-up procedure, repetitions, percentiles, failures, and quality results.
Conclusion
The SPEED blueprint is a useful map, but production inference is an experiment, not a checklist. Quantization changes the memory and quality trade-off. Paged KV changes allocation behavior. Continuous batching changes scheduling. Speculative decoding changes the decode loop. Each change must be measured against a fixed baseline and protected by quality, privacy, and rollback controls.
Use the ROUTE framework when deciding which model should handle a request, then use SPEED to serve that model efficiently. The strongest optimization is the one that improves a defined user-visible SLO without hiding regressions in quality or safety.
Trusted references
- vLLM: Speculative Decoding.
- vLLM: Paged Attention and current implementation note.
- vLLM: Quantized KV Cache.
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention.
- Leviathan et al., Fast Inference from Transformers via Speculative Decoding.
- Hugging Face Transformers Quantization Documentation.
- NVIDIA TensorRT-LLM Documentation.
- PromptSphere: LLM Model Routing and the ROUTE Framework.
Sources support the methods and measurement principles. They do not validate universal speedup percentages for every model or workload. Replace local benchmark placeholders with measurements from your own target stack before making performance claims.
Join the conversation