LLM Inference Optimization: Quantization, KV Cache & Batching

A practical guide to LLM inference optimization with quantization, KV cache, continuous batching, speculative decoding, benchmarks, and trade-offs.
SPEED framework for LLM inference optimization: quantization, KV cache, batching, speculative decoding, and deployment
The SPEED framework connects memory efficiency, scheduling, decoding, and deployment decisions.
Important: There is no universal percentage improvement for LLM inference optimization. The real result depends on the model, GPU, runtime version, context length, concurrency, sampling settings, and traffic mix. This article provides a reproducible test plan instead of inventing benchmark numbers.

Quick answer

To optimize a production LLM, first measure a fixed-model baseline. Then reduce memory pressure with an appropriate weight or KV-cache format, improve scheduling with continuous batching, test speculative decoding only on workloads where draft-token acceptance is healthy, and choose local or cloud deployment according to privacy, cost, reliability, and SLO requirements.

The useful sequence is measure → change one variable → measure again → check quality → canary → roll back if necessary. A smaller model or lower-precision model is not automatically faster if its runtime lacks an efficient kernel or if it causes retries and quality failures.

The SPEED framework

S — Shrink the mathUse INT8, FP8, or INT4 only after evaluating quality and kernel support on the target hardware.
P — Page and persistManage KV memory in blocks, reuse safe prefixes, and isolate cache entries by tenant and permissions.
E — Exploit demandUse in-flight batching and token budgets, but measure p95 latency rather than GPU utilization alone.
E — Early draftingUse a draft model or another proposer only when acceptance and memory conditions justify it.
D — Deploy deliberatelyCompare local, cloud, and managed inference using total cost, privacy, reliability, and operational effort.

This framework complements the ROUTE framework for model routing. Routing decides which model should answer; inference optimization decides how that model is served.

S — Quantization that respects quality

Quantization represents weights, activations, or cached keys and values with fewer bits or a lower-precision floating-point format. It can reduce VRAM use and sometimes improve throughput, but the result is workload-dependent. A useful deployment decision must include both system metrics and task quality.

What should be quantized?

TargetPotential benefitMain riskValidation
WeightsLower model memory and often lower bandwidth pressureQuality loss, especially on sensitive layers or tasksTask accuracy, structured-output validity, VRAM
ActivationsLower memory traffic in supported kernelsRuntime and calibration incompatibilityPrefill latency and end-to-end quality
KV cacheMore context tokens per GPUAttention-quality loss or backend limitationsLong-context quality, OOM rate, TPOT

INT4 AWQ or GPTQ can be valuable for some decoder-only models, while INT8 or FP8 may provide a safer quality/performance trade-off. Do not describe INT8 as universally safe or INT4 as universally faster. The implementation, group size, calibration data, and hardware kernel matter.

For current deployment details, consult the Hugging Face quantization documentation and TensorRT-LLM documentation. Record the exact format and runtime version in every benchmark.

P — Page and persist the KV cache

Paged KV cache and LLM serving workflow for context memory management
Paged KV storage reduces allocation waste; it does not remove the cost of reading the relevant context during attention.

During autoregressive decoding, the model reuses keys and values from previous tokens. A paged cache stores those values in fixed-size blocks and maps logical sequences to physical memory. This reduces fragmentation and makes multi-request serving easier to schedule. The original PagedAttention paper describes the foundation, while the current vLLM documentation notes that its historical kernel explanation does not describe every detail of the current implementation.

Prefix reuse needs privacy boundaries

Prefix caching can reuse a shared system prompt or a stable document prefix, but the cache key must include the model version, tokenizer or template version, tenant, permissions, and relevant policy state. Never share a cached prefix across users merely because the text looks similar. Add TTLs, invalidation rules, and access checks.

KV quantization and offload

vLLM documents FP8 KV-cache options with calibration and layer-skipping controls. The benefit is more tokens in memory, not a guaranteed reduction in every latency metric. CPU offload can prevent an out-of-memory failure, but PCIe transfer and cache thrashing can make a request slower. Disk-backed cache should be treated as an exceptional capacity mechanism, not a default decode optimization.

E — Continuous batching and latency discipline

Continuous or in-flight batching admits and retires requests during the decode loop. It can improve aggregate utilization and throughput when requests have different lengths, but it can also increase queueing or tail latency if token budgets and scheduling priorities are poorly chosen.

MetricMeaningWhy it matters
TTFTTime to first tokenPrompt processing and queueing experience
TPOT / ITLTime per output token or inter-token latencyStreaming smoothness after the first token
ThroughputOutput tokens per second across requestsCapacity and unit economics
p95 / p99Tail latency percentilesReliability for the slowest users
GoodputThroughput under explicit SLO limitsPrevents fast-but-unusable benchmarks

Set a maximum number of new tokens, separate interactive and batch queues where appropriate, and report concurrency. A single tokens-per-second number without prompt length, output length, and concurrency is not a meaningful production comparison.

E — Speculative decoding: test acceptance, not hype

Speculative decoding uses a smaller proposer to draft several tokens and a larger target model to verify them. The original research describes a lossless sampling method under its assumptions, but the speedup depends on draft cost, token acceptance, target-model bottlenecks, sampling settings, and the serving workload.

The current vLLM guidance focuses on reducing inter-token latency in medium-to-low-QPS, memory-bound workloads and lists different proposer methods. Therefore, a fixed recommendation such as “always pair a 2–4B draft with a 7–13B target” is too broad. Test the pair on your own prompts.

Decision rule: keep speculative decoding only if it improves the target SLO after including draft-model memory, proposer compute, acceptance rate, and any change in throughput. If it increases resource use without improving user-visible latency, disable it.

Reproducible experiments you can run

The following experiments are designed to produce real measurements on your hardware. They intentionally do not claim a universal result. Publish the output with the GPU, driver, CUDA version, model revision, runtime version, prompt set, concurrency, and sampling configuration.

Experiment 1: establish a baseline

# Example vLLM server; pin the versions in your own environment
vllm serve MODEL_ID \
  --dtype auto \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.90

# Run the official benchmark tool available in your installed vLLM version
python benchmarks/benchmark_serving.py \
  --backend vllm \
  --model MODEL_ID \
  --dataset-name random \
  --random-input-len 1024 \
  --random-output-len 256 \
  --num-prompts 200 \
  --request-rate inf

Record TTFT, TPOT, output tokens per second, p50, p95, p99, GPU memory, and request failures. Repeat at concurrency 1, 4, 16, and the expected production level. Store the command and raw output beside the published table.

Experiment 2: compare precision without hiding quality loss

# Run the same prompt set against baseline and quantized variants
python evaluate_prompts.py \
  --model baseline \
  --prompts prompts.jsonl \
  --seed 42 \
  --output baseline.jsonl

python evaluate_prompts.py \
  --model quantized \
  --prompts prompts.jsonl \
  --seed 42 \
  --output quantized.jsonl

Your prompt set should include extraction, reasoning, RAG, tool-use, refusal, and long-context cases. Compare exact match, JSON validity, citation coverage, task-specific score, and human review on failure slices. A lower perplexity score alone is not enough for a product decision.

Experiment 3: test speculative decoding

Keep the target model, prompt set, output limit, and sampling settings fixed. Change only the proposer configuration. Report proposer acceptance rate, TPOT, end-to-end latency, throughput, GPU memory, and quality equivalence. Test low-QPS and high-QPS separately because a latency win at low concurrency may not be a throughput win under a full queue.

Do publishHardware, versions, prompts, concurrency, raw metrics, quality results, and failed configurations.
Do not publishA single best-case tokens/s number without workload details or an unsupported percentage guarantee.

D — Deployment roadmap

  1. Baseline first. Define SLOs for TTFT, TPOT, p95 latency, throughput, quality, cost, and error rate before changing the stack.
  2. Reduce memory pressure. Test weight precision and KV-cache settings separately. Keep a rollback configuration and a quality canary set.
  3. Improve scheduling. Add continuous batching, token budgets, queue priorities, and observability. Watch p95 and queue time, not only GPU utilization.
  4. Test speculative decoding. Measure acceptance and target SLOs at the traffic levels that matter. Remove it if the total system becomes slower or more expensive.
  5. Choose deployment deliberately. Compare local GPUs, cloud GPUs, and managed inference using total cost of ownership, data residency, reliability, and operational effort.

Inference optimization also connects to AI observability, agent memory design, and zero-trust access control. A faster server that leaks data or cannot explain a regression is not production-ready.

Common failure modes

FailureWhy it happensControl
INT4 looks fast but quality dropsCalibration or kernel does not fit the taskSlice-level quality canary and rollback
Batching raises p99Long requests starve interactive trafficToken caps, queue separation, priority limits
Speculation adds overheadLow draft acceptance or high QPSMeasure acceptance and disable by policy
Prefix cache leaks contextCache key ignores tenant or permissionsIsolation, ACL checks, TTL, invalidation
Benchmark cannot be reproducedMissing versions and workload detailsPublish config, prompts, raw output, and seed

FAQ

What is the fastest optimization to try?

Start with measurement and a memory profile. If the model fits comfortably and the workload is low-concurrency, quantization may be more useful than speculative decoding. If the GPU is underfilled by many concurrent requests, scheduling may have the larger effect.

Does quantization always reduce cost?

No. It can reduce memory and allow more requests per GPU, but a slower kernel, quality failures, retries, or a more expensive validation path can cancel the saving.

Should I use CPU or disk offload?

Use offload as a capacity and resilience option after measuring transfer cost. It is not a free latency optimization.

How do I make the article's benchmark credible?

Publish the exact model revision, runtime and driver versions, hardware, prompt distribution, input and output lengths, concurrency, sampling settings, warm-up procedure, repetitions, percentiles, failures, and quality results.

Conclusion

The SPEED blueprint is a useful map, but production inference is an experiment, not a checklist. Quantization changes the memory and quality trade-off. Paged KV changes allocation behavior. Continuous batching changes scheduling. Speculative decoding changes the decode loop. Each change must be measured against a fixed baseline and protected by quality, privacy, and rollback controls.

Use the ROUTE framework when deciding which model should handle a request, then use SPEED to serve that model efficiently. The strongest optimization is the one that improves a defined user-visible SLO without hiding regressions in quality or safety.

Trusted references

  1. vLLM: Speculative Decoding.
  2. vLLM: Paged Attention and current implementation note.
  3. vLLM: Quantized KV Cache.
  4. Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention.
  5. Leviathan et al., Fast Inference from Transformers via Speculative Decoding.
  6. Hugging Face Transformers Quantization Documentation.
  7. NVIDIA TensorRT-LLM Documentation.
  8. PromptSphere: LLM Model Routing and the ROUTE Framework.

Sources support the methods and measurement principles. They do not validate universal speedup percentages for every model or workload. Replace local benchmark placeholders with measurements from your own target stack before making performance claims.

FE
Written and reviewed by Fouad El Mourabit

Fouad El Mourabit is a Morocco-based technology writer and editor covering artificial intelligence, search, content systems, software, and practical digital workflows.

For corrections or updated sources, visit the PromptSphere About page.

PromptSphere Welcome to WhatsApp chat
Howdy! How can we help you today?
Type here...