Infrastructure5 min read

GPU memory for LLM inference: how to size and choose a GPU

The first question in any LLM deployment is whether the model fits, and the second is how many users fit alongside it. Both come down to GPU memory. This guide gives you the formulas, a small calculator, and a way to compare common data center GPUs before you rent one.

What uses GPU memory during inference

An inference server divides GPU memory between three things:

  • Model weights. Fixed for a given model and precision, and loaded once at startup.
  • KV cache. Grows with the number of tokens each request holds and the number of requests running at once. This is what limits concurrency.
  • Runtime overhead. Activations for the current step, CUDA graphs, and libraries. It is smaller, but it is not zero.

Engines such as vLLM reserve a fixed share of the GPU up front, load the weights, and give most of what remains to the KV cache. In vLLM v0.30.0 that share is set by --gpu-memory-utilization, which defaults to 0.92. vLLM engine arguments.

How much memory the weights need

Weight memory is the parameter count multiplied by the bytes per parameter. BF16 and FP16 use 2 bytes, FP8 and INT8 use 1, and 4-bit formats use about half a byte plus a small amount of scaling data.

Approximate weight memory in decimal gigabytes. Quantized formats add a little for scales.
Model sizeBF16 (2 bytes)FP8 (1 byte)4-bit (~0.5 bytes)
8 billion parameters16 GB8 GB~4 GB
32 billion parameters64 GB32 GB~16 GB
70 billion parameters140 GB70 GB~35 GB

Fitting the weights is only the start. A 70B model in BF16 needs more than one 80 GB GPU before any request runs, and a model that barely fits leaves no room for users. LLM quantization explained covers the formats and their accuracy tradeoffs.

How much memory the KV cache needs

For standard attention, every token a request holds stores one key and one value vector per layer and per KV head:

KV cache size
KV bytes = 2 × layers × KV heads × head dim × bytes per value × tokens × requests

Qwen3-8B has 36 layers, 8 KV heads, and a head dimension of 128 in its published configuration. In BF16 that is 144 KiB per token, or 576 MiB for one 4,096-token conversation. Qwen3-32B has 64 layers with the same head layout, so it needs 256 KiB per token. The KV cache explained derives the formula and includes an interactive calculator.

Memory and bandwidth of common inference GPUs

Figures from NVIDIA’s product pages. Capacity decides what fits; bandwidth decides how fast each answer streams.
GPUMemoryMemory bandwidth
L40S48 GB GDDR6864 GB/s
A100 80GB SXM80 GB HBM2e2,039 GB/s
H100 SXM80 GB3.35 TB/s
H100 NVL94 GB3.9 TB/s
H200141 GB4.8 TB/s

The A100 80GB and H100 SXM hold the same amount, so they fit the same models and batch sizes. The H100 reads that memory about 1.6 times faster and supports FP8 in hardware, which makes each decode step quicker.

Estimate how many requests fit on a GPU

This calculator subtracts the weights and a 2 GB runtime reserve from 92 percent of each GPU’s memory, then counts how many 4,096-token conversations the remaining KV cache can hold. The reserve is a planning assumption; your engine reports the real figure at startup.

Run with Python 3; no GPU or packages required
# Rough planning estimate: weights + KV cache for one model on one GPU.
GB = 1e9

models = {  # parameters, layers, KV heads, head dim (from each config.json)
    "Qwen3-8B": (8.19e9, 36, 8, 128),
    "Qwen3-32B": (32.76e9, 64, 8, 128),
}
gpus = {"L40S": 48, "H100 SXM": 80, "H200": 141}  # GB, from NVIDIA spec pages


def kv_bytes_per_token(layers, kv_heads, head_dim, bytes_per_value=2):
    return 2 * layers * kv_heads * head_dim * bytes_per_value  # K and V


def sequences_that_fit(gpu_gb, model, weight_bytes=2, context=4096,
                       utilization=0.92, reserve_gb=2):
    params, layers, kv_heads, head_dim = model
    weights = params * weight_bytes
    budget = gpu_gb * GB * utilization - reserve_gb * GB - weights
    per_sequence = kv_bytes_per_token(layers, kv_heads, head_dim) * context
    return max(0, int(budget // per_sequence)), weights / GB


for name, model in models.items():
    kib = kv_bytes_per_token(*model[1:]) / 1024
    print(f"{name}: {kib:.0f} KiB of KV cache per token")
    for gpu, memory in gpus.items():
        for label, weight_bytes in (("BF16", 2), ("FP8", 1)):
            fit, weights = sequences_that_fit(memory, model, weight_bytes)
            print(f"  {gpu:<10} {label}: weights {weights:5.1f} GB, "
                  f"{fit:>3} sequences of 4,096 tokens")
Calculated output
Qwen3-8B: 144 KiB of KV cache per token
  L40S       BF16: weights  16.4 GB,  42 sequences of 4,096 tokens
  L40S       FP8: weights   8.2 GB,  56 sequences of 4,096 tokens
  H100 SXM   BF16: weights  16.4 GB,  91 sequences of 4,096 tokens
  H100 SXM   FP8: weights   8.2 GB, 104 sequences of 4,096 tokens
  H200       BF16: weights  16.4 GB, 184 sequences of 4,096 tokens
  H200       FP8: weights   8.2 GB, 197 sequences of 4,096 tokens
Qwen3-32B: 256 KiB of KV cache per token
  L40S       BF16: weights  65.5 GB,   0 sequences of 4,096 tokens
  L40S       FP8: weights  32.8 GB,   8 sequences of 4,096 tokens
  H100 SXM   BF16: weights  65.5 GB,   5 sequences of 4,096 tokens
  H100 SXM   FP8: weights  32.8 GB,  36 sequences of 4,096 tokens
  H200       BF16: weights  65.5 GB,  57 sequences of 4,096 tokens
  H200       FP8: weights  32.8 GB,  88 sequences of 4,096 tokens

The 8B model fits comfortably everywhere, so the choice between GPUs is about speed and price. The 32B model is the interesting case. In BF16 it does not fit on an L40S at all and leaves an H100 with room for only five long conversations. FP8 weights multiply that by about seven, and the H200’s extra memory changes the picture entirely. These are estimates: real counts also depend on prefix sharing, actual conversation lengths, and how the engine allocates cache blocks.

What to do when the model does not fit

  • Shorten the maximum context. If users never send 32,000 tokens, do not reserve room for them. vLLM’s --max-model-len caps it. vLLM: conserving memory.
  • Quantize the weights to FP8 or 4-bit, after checking accuracy on your own prompts.
  • Quantize the KV cache. An FP8 cache halves its size. vLLM: quantized KV cache.
  • Use a GPU with more memory, such as moving from an H100 to an H200.
  • Split the model across GPUs with tensor parallelism. The vLLM guide recommends a single GPU when the model fits, tensor parallelism across the GPUs of one node when it does not, and adding pipeline parallelism across nodes only beyond that. vLLM: parallelism and scaling.

GPU memory FAQ

How much GPU memory do I need for a 7B or 8B model? About 16 GB for BF16 weights, plus room for the KV cache. A 24 GB card runs it for a few users; a 48 GB or 80 GB card serves many more.

Can I run a 70B model on one GPU? Only if it is quantized. In BF16 the weights alone are about 140 GB, so you need several 80 GB GPUs with tensor parallelism, or 4-bit weights on a single large GPU with limited cache room.

Is more memory or more bandwidth better? Memory decides what fits and how many users share a GPU; bandwidth decides how fast each answer streams. How to learn inference engineering shows how to estimate the bandwidth limit.

Which cloud setup should I use? See AI cloud infrastructure for LLM inference for deployment options, scaling, and cost per million tokens. The model directory lists architectures and engine support for many open models.

Sources and further reading

Primary documentation and research behind this guide.

  1. vLLM v0.30.0: engine arguments
  2. Qwen3-8B configuration
  3. Qwen3-32B configuration
  4. NVIDIA L40S specifications
  5. NVIDIA A100 specifications
  6. NVIDIA H100 specifications
  7. NVIDIA H200 specifications
  8. vLLM v0.30.0: conserving memory
  9. vLLM: quantized KV cache
  10. vLLM: parallelism and scaling

Keep learning

Related guideAI cloud infrastructure for LLM inference: a practical guideRelated guideWhat is AI inference? How trained models answer in productionRelated guideLLM quantization explained: FP8, INT8, INT4, AWQ, GPTQRelated guideKV cache explained: the formula, a diagram, and a memory exampleAcademy courseInference Engineering FoundationsAcademy coursevLLM
Back to all articles