FlashAttention vs PagedAttention at a glance
| Question | FlashAttention | PagedAttention |
|---|---|---|
| Main problem | Moving large attention intermediates through GPU memory | Wasted or duplicated storage for growing KV caches |
| Core idea | Tile attention computation and reduce memory traffic | Map logical token blocks to physical cache blocks |
| Useful outcome | More efficient attention execution | More usable cache capacity for concurrent requests |
| What it does not do | Make model weights smaller | Make model weights smaller or remove attention computation |
| Can they work together? | Yes, a compatible kernel can read a paged cache | Yes, paging can be paired with efficient attention kernels |
The original FlashAttention paper focuses on attention’s data movement. The PagedAttention paper focuses on dynamically managing KV state in a serving system. They are useful at different layers of the same request path.
What does FlashAttention optimize?
Standard attention computes softmax(Q × Kᵀ / √d) × V. A straightforward implementation writes a large matrix of scores, then probabilities, to GPU memory. FlashAttention tiles the calculation, uses fast on-chip storage, and accumulates the result without materializing the entire matrix in high-bandwidth memory. It computes exact attention in the algorithmic sense, with normal floating-point differences. FlashAttention research.
For one sequence with 32 query heads, an 8,192-token dense score tensor in a two-byte dtype occupies 4 GiB. At 16,384 tokens it occupies 16 GiB. Those are calculated sizes for one hypothetical full tensor, not measured peak usage: masking, fused kernels, intermediate precision, and the implementation change actual storage.
Avoiding that tensor does not make dense attention’s pairwise arithmetic linear. It reduces a major data-movement cost. It also does not eliminate the persistent KV state needed during generation. Follow the foundations course for the attention equations and the decode loop.
What does PagedAttention optimize?
An LLM server holds a growing KV cache for each active request. Reserving a large contiguous region for every possible output wastes capacity. PagedAttention divides the cache into fixed-size blocks and uses a block table to locate them, so a request can grow without requiring contiguous physical storage. Compatible blocks can also be shared. PagedAttention research.
Paging changes allocation, not the number of values required to represent an unshared token. There is still space overhead in a partially filled final block, plus metadata. More usable cache capacity can let a scheduler admit a larger batch; whether that improves user latency depends on the traffic and scheduler. See continuous batching explained.
Calculate attention intermediates and paged cache storage
This Python example keeps the two memory questions separate. First it estimates one dense score tensor. Then it compares reserving 128 tokens per request with allocating sixteen-token cache blocks for three invented request lengths. All numbers are illustrative; no GPU or model is loaded.
# One full score tensor: batch 1, 32 query heads, 2 bytes per element.
for tokens in (4096, 8192, 16384):
score_bytes = 32 * tokens * tokens * 2
print(f"{tokens:,} tokens: dense scores {score_bytes / 2**30:.1f} GiB")
# Separate example: KV tensors for active requests, without block sharing.
lengths = [19, 45, 70]
block_size = 16
reserved_per_request = 128
kv_bytes_per_token = 128 * 1024 # Chosen teaching assumption, all layers.
reserved_tokens = len(lengths) * reserved_per_request
paged_tokens = sum(
((length + block_size - 1) // block_size) * block_size
for length in lengths
)
used_tokens = sum(lengths)
for label, tokens in (("Fixed reservation", reserved_tokens),
("Paged allocation", paged_tokens),
("Used KV tensors", used_tokens)):
mib = tokens * kv_bytes_per_token / 2**20
print(f"{label}: {tokens} token slots, {mib:.2f} MiB")
print(f"Unused paged slots: {paged_tokens - used_tokens}")4,096 tokens: dense scores 1.0 GiB
8,192 tokens: dense scores 4.0 GiB
16,384 tokens: dense scores 16.0 GiB
Fixed reservation: 384 token slots, 48.00 MiB
Paged allocation: 160 token slots, 20.00 MiB
Used KV tensors: 134 token slots, 16.75 MiB
Unused paged slots: 26The paged version rounds 19 tokens to 32 slots, 45 to 48, and 70 to 80. It allocates 20 MiB rather than 48 MiB while keeping the same 134 tokens of state. That saves reserved capacity in this example; it is not a speed benchmark. The two parts measure different objects, so do not subtract the score-tensor estimate from the KV-cache budget.
To size an actual server, use its layer count, KV heads, head dimension, cache dtype, and active lengths. Query heads and KV heads can differ with grouped-query attention. GPU memory for LLM inference covers weights, KV state, and runtime headroom together.
Can FlashAttention and PagedAttention work together?
Yes. The FlashAttention project documents support for a paged KV cache through a block_table in its cache-aware interface. This is direct evidence that efficient attention execution and noncontiguous cache storage can be combined. Support still depends on the specific interface and backend. FlashAttention implementation.
An inference engine chooses an attention backend that works with the model, device, dtype, and attention features. vLLM publishes a backend support matrix covering constraints such as head sizes, cache dtypes, sliding windows, and attention variants. Check that matrix for your release before forcing a backend.
How prefill and decode change the diagnosis
Prefill processes many prompt positions, while a decode step usually adds one position per active request. That changes the shapes of the attention work. The FlashAttention repository includes a cache-aware inference interface for unequal query and cache lengths, so describing FlashAttention as only a training technique or only a prefill technique is misleading. FlashAttention cache interface.
| Observed symptom | What to inspect | A useful next experiment |
|---|---|---|
| Long prompts start slowly | Prefill time, selected kernel, and queueing | Compare prompt-length buckets at the same request rate |
| Few concurrent requests fit | Weight footprint, KV budget, active lengths, and preemption | Estimate cache capacity and compare it with engine logs |
| Repeated long context is expensive | Shared token prefixes and cache reuse | Compare cold and warm runs with prefix caching |
| Answers stream slowly | Decode time, batch size, memory traffic, and communication | Profile the serving step before changing a kernel |
Treat these as hypotheses because the same symptom can have several causes. LLM inference metrics explains how to separate time to first token from inter-token latency. Prefix caching explained covers the repeated-context experiment.
FlashAttention and PagedAttention FAQ
Which one should I enable first? Start from the serving engine’s supported defaults, record its backend, and establish a workload baseline. Use a measured bottleneck to decide what to change.
Are their published speedups directly comparable? No. Kernel timing, training time, and serving throughput describe different experiments. Match workload, hardware, quality, and latency constraints before drawing a conclusion.
Will they solve a model-weight out-of-memory error? Neither changes the weight tensor size. First budget the weights and runtime state, then consider quantization or model parallelism.
Where can I study the implementation? Use the vLLM course for cache allocation and scheduling, and the PyTorch course for tensors and attention operations.
Sources and further reading
Primary documentation and research behind this guide.