What is prefix caching in LLM inference?
Prefix caching retains the keys and values computed for the beginning of an earlier prompt. When a later request starts with the same token sequence, the server can reuse that state and process the uncached remainder. vLLM calls this automatic prefix caching, or APC. Its main direct benefit is less prefill work and a shorter time to first token; it does not remove the token-by-token work of generating an answer. vLLM automatic prefix caching.
Consider a document assistant whose prompt contains a handbook followed by a question. Reading the handbook once can prepare reusable state for later questions. The server still reads each new question and generates a fresh response. Start with LLM inference explained if prefill and decode are new terms.
Prefix caching vs KV caching vs response caching
| Technique | What is reused | Example |
|---|---|---|
| KV cache during generation | Attention state for tokens already processed in a request | Generate the next word without rebuilding earlier keys and values |
| Prefix caching across requests | Compatible KV state for a shared prompt beginning | Ask different questions about the same handbook |
| Response caching in the application | A previously generated answer | Return the stored answer to an identical lookup |
A response cache needs application rules for freshness, permissions, and when two questions count as equivalent. A prefix cache leaves response generation to the model. Reusing state also does not guarantee identical sampled text across runs. The KV cache explained covers the underlying tensors and their memory cost.
Why matching words in the middle are not enough
For a causal transformer, a token’s state depends on the tokens before it. Identical text after a different beginning is therefore not interchangeable cached state. vLLM identifies reusable blocks using their tokens, the preceding block’s hash, and relevant extra identifiers such as adapters. Its documented design caches full blocks and evicts cached blocks as space is needed. vLLM prefix cache design.
More reusable:
[fixed instructions] [fixed handbook] [new question]
Less reusable:
[new request ID] [fixed instructions] [fixed handbook] [new question]This suggests a prompt-design experiment: move a changing request ID out of the model prompt and keep it in request metadata. Compare the rendered token prefixes before and after. Only make this change if the ID is unnecessary to the task; cacheability should not change the instructions the model needs.
A small simulation of prefix cache hits
This standalone Python example uses invented token IDs and four-token blocks. The block size is a teaching choice, not a vLLM default. Each cache key includes the entire prefix through a block boundary, so changing the first token invalidates every later match.
block_size = 4
stable = list(range(8))
requests = [
("cold", stable + list(range(8, 12))),
("warm", stable + list(range(20, 24))),
("changed first token", [99] + stable[1:] + list(range(30, 34))),
("warm again", stable + list(range(40, 44))),
]
cached_prefixes = set()
total_tokens = total_computed = 0
for label, tokens in requests:
reused = 0
boundaries = range(block_size, len(tokens) + 1, block_size)
for end in boundaries:
if tuple(tokens[:end]) not in cached_prefixes:
break
reused = end
computed = len(tokens) - reused
for end in boundaries:
cached_prefixes.add(tuple(tokens[:end]))
total_tokens += len(tokens)
total_computed += computed
print(f"{label}: reused {reused}, computed {computed}")
saved = total_tokens - total_computed
print(f"Total: {total_computed}/{total_tokens} tokens computed; "
f"{saved / total_tokens:.1%} avoided")cold: reused 0, computed 12
warm: reused 8, computed 4
changed first token: reused 0, computed 12
warm again: reused 8, computed 4
Total: 32/48 tokens computed; 33.3% avoidedThe third request loses all reuse despite sharing seven of the eight instruction tokens. The fourth can still find the original prefix. This counts token positions, not FLOPs or milliseconds: it omits eviction, lookup overhead, scheduling, model-specific state, and any final prompt-token recomputation. A 33.3 percent reduction here is not a predicted latency improvement.
Design cacheable prompts without changing the task
- Inspect the actual input. Compare token IDs after applying the chat template, rather than comparing only the visible user text.
- Keep a stable opening. Use consistent instructions and tool definitions where their meaning stays the same. Put changing questions later when that preserves the task.
- Make reference context reproducible. For a document assistant, keep document order and formatting stable. Do not reorder retrieval results just for caching if relevance order matters.
- Version prompt changes. Record the model revision, template, adapter, and prompt version alongside measurements so a drop in reuse has an explanation.
Treat these as application experiments. A support bot with a long fixed policy and short questions is a promising candidate. A service with mostly unrelated prompts may have little reusable work. Long answers can also dominate total latency even when their prompts are highly cacheable.
How to benchmark prefix caching in vLLM
Use the vLLM tutorial to establish a working server first. The documented offline setting is enable_prefix_caching=True; verify the configuration for your installed release rather than assuming its defaults. Repeated document questions and multi-turn conversations are the workloads highlighted in the vLLM APC guide.
- Build three prompt groups: shared long prefixes, unique prefixes of comparable length, and your normal traffic mix. Keep output limits consistent.
- Separate cold and warm runs: record the initial request, then repeated questions after warming the prefix. Keep startup and model-loading time out of request latency.
- Compare caching enabled and disabled: use the same model, hardware, arrival rate, and input set, and reset cache state between configurations.
- Report user-facing metrics: median and p95 time to first token, end-to-end latency, successful requests per second, and the fraction meeting your latency target.
Also repeat under realistic concurrency, where queueing and cache pressure can change the result. vLLM’s optimization guide explains the tradeoffs around scheduling and memory. Our inference metrics guide defines the latency and goodput measures. The simulation above is not a GPU benchmark.
Keep cache reuse inside the right trust boundary
A shared prefix cache can expose whether a prompt has been processed through timing differences. vLLM documents cache_salt to separate reuse between trust groups, and warns about collision risks with noncryptographic cache hashes. vLLM cache isolation.
For a multi-tenant application, have the authenticated backend choose the isolation value. Do not let an untrusted caller select another tenant’s cache group. Cache isolation complements authorization: it does not grant permission to retrieve a document or return a stored answer.
Prefix caching FAQ
Will it make every response faster? Measure the whole request. Short prompts, long outputs, cache misses, or a busy queue can make the benefit small.
Should I use semantic response caching instead? That is a separate application decision. Before reusing an answer, define how you will preserve factual freshness, permissions, and task-specific accuracy.
What should I learn next? Study FlashAttention vs PagedAttention to connect prompt reuse with attention kernels and cache allocation. The vLLM course walks through the serving system in more depth.
Sources and further reading
Primary documentation and research behind this guide.