What uses GPU memory during inference
An inference server divides GPU memory between three things:
- Model weights. Fixed for a given model and precision, and loaded once at startup.
- KV cache. Grows with the number of tokens each request holds and the number of requests running at once. This is what limits concurrency.
- Runtime overhead. Activations for the current step, CUDA graphs, and libraries. It is smaller, but it is not zero.
Engines such as vLLM reserve a fixed share of the GPU up front, load the weights, and give most of what remains to the KV cache. In vLLM v0.30.0 that share is set by --gpu-memory-utilization, which defaults to 0.92. vLLM engine arguments.
How much memory the weights need
Weight memory is the parameter count multiplied by the bytes per parameter. BF16 and FP16 use 2 bytes, FP8 and INT8 use 1, and 4-bit formats use about half a byte plus a small amount of scaling data.
| Model size | BF16 (2 bytes) | FP8 (1 byte) | 4-bit (~0.5 bytes) |
|---|---|---|---|
| 8 billion parameters | 16 GB | 8 GB | ~4 GB |
| 32 billion parameters | 64 GB | 32 GB | ~16 GB |
| 70 billion parameters | 140 GB | 70 GB | ~35 GB |
Fitting the weights is only the start. A 70B model in BF16 needs more than one 80 GB GPU before any request runs, and a model that barely fits leaves no room for users. LLM quantization explained covers the formats and their accuracy tradeoffs.
How much memory the KV cache needs
For standard attention, every token a request holds stores one key and one value vector per layer and per KV head:
KV bytes = 2 × layers × KV heads × head dim × bytes per value × tokens × requestsQwen3-8B has 36 layers, 8 KV heads, and a head dimension of 128 in its published configuration. In BF16 that is 144 KiB per token, or 576 MiB for one 4,096-token conversation. Qwen3-32B has 64 layers with the same head layout, so it needs 256 KiB per token. The KV cache explained derives the formula and includes an interactive calculator.
Memory and bandwidth of common inference GPUs
| GPU | Memory | Memory bandwidth |
|---|---|---|
| L40S | 48 GB GDDR6 | 864 GB/s |
| A100 80GB SXM | 80 GB HBM2e | 2,039 GB/s |
| H100 SXM | 80 GB | 3.35 TB/s |
| H100 NVL | 94 GB | 3.9 TB/s |
| H200 | 141 GB | 4.8 TB/s |
The A100 80GB and H100 SXM hold the same amount, so they fit the same models and batch sizes. The H100 reads that memory about 1.6 times faster and supports FP8 in hardware, which makes each decode step quicker.
Estimate how many requests fit on a GPU
This calculator subtracts the weights and a 2 GB runtime reserve from 92 percent of each GPU’s memory, then counts how many 4,096-token conversations the remaining KV cache can hold. The reserve is a planning assumption; your engine reports the real figure at startup.
# Rough planning estimate: weights + KV cache for one model on one GPU.
GB = 1e9
models = { # parameters, layers, KV heads, head dim (from each config.json)
"Qwen3-8B": (8.19e9, 36, 8, 128),
"Qwen3-32B": (32.76e9, 64, 8, 128),
}
gpus = {"L40S": 48, "H100 SXM": 80, "H200": 141} # GB, from NVIDIA spec pages
def kv_bytes_per_token(layers, kv_heads, head_dim, bytes_per_value=2):
return 2 * layers * kv_heads * head_dim * bytes_per_value # K and V
def sequences_that_fit(gpu_gb, model, weight_bytes=2, context=4096,
utilization=0.92, reserve_gb=2):
params, layers, kv_heads, head_dim = model
weights = params * weight_bytes
budget = gpu_gb * GB * utilization - reserve_gb * GB - weights
per_sequence = kv_bytes_per_token(layers, kv_heads, head_dim) * context
return max(0, int(budget // per_sequence)), weights / GB
for name, model in models.items():
kib = kv_bytes_per_token(*model[1:]) / 1024
print(f"{name}: {kib:.0f} KiB of KV cache per token")
for gpu, memory in gpus.items():
for label, weight_bytes in (("BF16", 2), ("FP8", 1)):
fit, weights = sequences_that_fit(memory, model, weight_bytes)
print(f" {gpu:<10} {label}: weights {weights:5.1f} GB, "
f"{fit:>3} sequences of 4,096 tokens")Qwen3-8B: 144 KiB of KV cache per token
L40S BF16: weights 16.4 GB, 42 sequences of 4,096 tokens
L40S FP8: weights 8.2 GB, 56 sequences of 4,096 tokens
H100 SXM BF16: weights 16.4 GB, 91 sequences of 4,096 tokens
H100 SXM FP8: weights 8.2 GB, 104 sequences of 4,096 tokens
H200 BF16: weights 16.4 GB, 184 sequences of 4,096 tokens
H200 FP8: weights 8.2 GB, 197 sequences of 4,096 tokens
Qwen3-32B: 256 KiB of KV cache per token
L40S BF16: weights 65.5 GB, 0 sequences of 4,096 tokens
L40S FP8: weights 32.8 GB, 8 sequences of 4,096 tokens
H100 SXM BF16: weights 65.5 GB, 5 sequences of 4,096 tokens
H100 SXM FP8: weights 32.8 GB, 36 sequences of 4,096 tokens
H200 BF16: weights 65.5 GB, 57 sequences of 4,096 tokens
H200 FP8: weights 32.8 GB, 88 sequences of 4,096 tokensThe 8B model fits comfortably everywhere, so the choice between GPUs is about speed and price. The 32B model is the interesting case. In BF16 it does not fit on an L40S at all and leaves an H100 with room for only five long conversations. FP8 weights multiply that by about seven, and the H200’s extra memory changes the picture entirely. These are estimates: real counts also depend on prefix sharing, actual conversation lengths, and how the engine allocates cache blocks.
What to do when the model does not fit
- Shorten the maximum context. If users never send 32,000 tokens, do not reserve room for them. vLLM’s
--max-model-lencaps it. vLLM: conserving memory. - Quantize the weights to FP8 or 4-bit, after checking accuracy on your own prompts.
- Quantize the KV cache. An FP8 cache halves its size. vLLM: quantized KV cache.
- Use a GPU with more memory, such as moving from an H100 to an H200.
- Split the model across GPUs with tensor parallelism. The vLLM guide recommends a single GPU when the model fits, tensor parallelism across the GPUs of one node when it does not, and adding pipeline parallelism across nodes only beyond that. vLLM: parallelism and scaling.
GPU memory FAQ
How much GPU memory do I need for a 7B or 8B model? About 16 GB for BF16 weights, plus room for the KV cache. A 24 GB card runs it for a few users; a 48 GB or 80 GB card serves many more.
Can I run a 70B model on one GPU? Only if it is quantized. In BF16 the weights alone are about 140 GB, so you need several 80 GB GPUs with tensor parallelism, or 4-bit weights on a single large GPU with limited cache room.
Is more memory or more bandwidth better? Memory decides what fits and how many users share a GPU; bandwidth decides how fast each answer streams. How to learn inference engineering shows how to estimate the bandwidth limit.
Which cloud setup should I use? See AI cloud infrastructure for LLM inference for deployment options, scaling, and cost per million tokens. The model directory lists architectures and engine support for many open models.
Sources and further reading
Primary documentation and research behind this guide.