Serving systems5 min read

Prefix caching explained: faster LLM prompts with KV reuse

An assistant may read the same instructions or reference document thousands of times. Prefix caching lets an inference server reuse that work. This guide explains what gets cached, shows why prompt order matters, and gives you a repeatable way to evaluate the benefit for your application.

What is prefix caching in LLM inference?

Prefix caching retains the keys and values computed for the beginning of an earlier prompt. When a later request starts with the same token sequence, the server can reuse that state and process the uncached remainder. vLLM calls this automatic prefix caching, or APC. Its main direct benefit is less prefill work and a shorter time to first token; it does not remove the token-by-token work of generating an answer. vLLM automatic prefix caching.

Consider a document assistant whose prompt contains a handbook followed by a question. Reading the handbook once can prepare reusable state for later questions. The server still reads each new question and generates a fresh response. Start with LLM inference explained if prefill and decode are new terms.

Prefix caching vs KV caching vs response caching

Three caches that reuse different kinds of work.
TechniqueWhat is reusedExample
KV cache during generationAttention state for tokens already processed in a requestGenerate the next word without rebuilding earlier keys and values
Prefix caching across requestsCompatible KV state for a shared prompt beginningAsk different questions about the same handbook
Response caching in the applicationA previously generated answerReturn the stored answer to an identical lookup

A response cache needs application rules for freshness, permissions, and when two questions count as equivalent. A prefix cache leaves response generation to the model. Reusing state also does not guarantee identical sampled text across runs. The KV cache explained covers the underlying tensors and their memory cost.

Why matching words in the middle are not enough

For a causal transformer, a token’s state depends on the tokens before it. Identical text after a different beginning is therefore not interchangeable cached state. vLLM identifies reusable blocks using their tokens, the preceding block’s hash, and relevant extra identifiers such as adapters. Its documented design caches full blocks and evicts cached blocks as space is needed. vLLM prefix cache design.

Put stable context before changing input
More reusable:
  [fixed instructions] [fixed handbook] [new question]

Less reusable:
  [new request ID] [fixed instructions] [fixed handbook] [new question]

This suggests a prompt-design experiment: move a changing request ID out of the model prompt and keep it in request metadata. Compare the rendered token prefixes before and after. Only make this change if the ID is unnecessary to the task; cacheability should not change the instructions the model needs.

A small simulation of prefix cache hits

This standalone Python example uses invented token IDs and four-token blocks. The block size is a teaching choice, not a vLLM default. Each cache key includes the entire prefix through a block boundary, so changing the first token invalidates every later match.

Run with Python 3; no GPU or packages required
block_size = 4
stable = list(range(8))
requests = [
    ("cold", stable + list(range(8, 12))),
    ("warm", stable + list(range(20, 24))),
    ("changed first token", [99] + stable[1:] + list(range(30, 34))),
    ("warm again", stable + list(range(40, 44))),
]
cached_prefixes = set()
total_tokens = total_computed = 0

for label, tokens in requests:
    reused = 0
    boundaries = range(block_size, len(tokens) + 1, block_size)
    for end in boundaries:
        if tuple(tokens[:end]) not in cached_prefixes:
            break
        reused = end
    computed = len(tokens) - reused
    for end in boundaries:
        cached_prefixes.add(tuple(tokens[:end]))
    total_tokens += len(tokens)
    total_computed += computed
    print(f"{label}: reused {reused}, computed {computed}")

saved = total_tokens - total_computed
print(f"Total: {total_computed}/{total_tokens} tokens computed; "
      f"{saved / total_tokens:.1%} avoided")
Calculated output
cold: reused 0, computed 12
warm: reused 8, computed 4
changed first token: reused 0, computed 12
warm again: reused 8, computed 4
Total: 32/48 tokens computed; 33.3% avoided

The third request loses all reuse despite sharing seven of the eight instruction tokens. The fourth can still find the original prefix. This counts token positions, not FLOPs or milliseconds: it omits eviction, lookup overhead, scheduling, model-specific state, and any final prompt-token recomputation. A 33.3 percent reduction here is not a predicted latency improvement.

Design cacheable prompts without changing the task

  1. Inspect the actual input. Compare token IDs after applying the chat template, rather than comparing only the visible user text.
  2. Keep a stable opening. Use consistent instructions and tool definitions where their meaning stays the same. Put changing questions later when that preserves the task.
  3. Make reference context reproducible. For a document assistant, keep document order and formatting stable. Do not reorder retrieval results just for caching if relevance order matters.
  4. Version prompt changes. Record the model revision, template, adapter, and prompt version alongside measurements so a drop in reuse has an explanation.

Treat these as application experiments. A support bot with a long fixed policy and short questions is a promising candidate. A service with mostly unrelated prompts may have little reusable work. Long answers can also dominate total latency even when their prompts are highly cacheable.

How to benchmark prefix caching in vLLM

Use the vLLM tutorial to establish a working server first. The documented offline setting is enable_prefix_caching=True; verify the configuration for your installed release rather than assuming its defaults. Repeated document questions and multi-turn conversations are the workloads highlighted in the vLLM APC guide.

  1. Build three prompt groups: shared long prefixes, unique prefixes of comparable length, and your normal traffic mix. Keep output limits consistent.
  2. Separate cold and warm runs: record the initial request, then repeated questions after warming the prefix. Keep startup and model-loading time out of request latency.
  3. Compare caching enabled and disabled: use the same model, hardware, arrival rate, and input set, and reset cache state between configurations.
  4. Report user-facing metrics: median and p95 time to first token, end-to-end latency, successful requests per second, and the fraction meeting your latency target.

Also repeat under realistic concurrency, where queueing and cache pressure can change the result. vLLM’s optimization guide explains the tradeoffs around scheduling and memory. Our inference metrics guide defines the latency and goodput measures. The simulation above is not a GPU benchmark.

Keep cache reuse inside the right trust boundary

A shared prefix cache can expose whether a prompt has been processed through timing differences. vLLM documents cache_salt to separate reuse between trust groups, and warns about collision risks with noncryptographic cache hashes. vLLM cache isolation.

For a multi-tenant application, have the authenticated backend choose the isolation value. Do not let an untrusted caller select another tenant’s cache group. Cache isolation complements authorization: it does not grant permission to retrieve a document or return a stored answer.

Prefix caching FAQ

Will it make every response faster? Measure the whole request. Short prompts, long outputs, cache misses, or a busy queue can make the benefit small.

Should I use semantic response caching instead? That is a separate application decision. Before reusing an answer, define how you will preserve factual freshness, permissions, and task-specific accuracy.

What should I learn next? Study FlashAttention vs PagedAttention to connect prompt reuse with attention kernels and cache allocation. The vLLM course walks through the serving system in more depth.

Sources and further reading

Primary documentation and research behind this guide.

  1. vLLM v0.30.0: automatic prefix caching
  2. vLLM v0.30.0: prefix cache design and isolation
  3. vLLM v0.30.0: optimization and tuning

Keep learning

Related guideFlashAttention vs PagedAttention: what each optimizesRelated guideLLM inference metrics: TTFT, TPOT, ITL, and throughputRelated guideKV cache explained: the formula, a diagram, and a memory exampleRelated guidevLLM tutorial: serve your first model with DockerAcademy courseInference Engineering FoundationsAcademy coursevLLMAcademy courseSGLang
Back to all articles