Attention and memory5 min read

FlashAttention vs PagedAttention: what each optimizes

FlashAttention and PagedAttention both make attention more efficient, but they address different sources of waste. Understanding the distinction helps you diagnose slow prompts, cache pressure, and limited concurrency without treating every serving problem as a kernel problem.

FlashAttention vs PagedAttention at a glance

The distinction is computation efficiency versus cache organization.
QuestionFlashAttentionPagedAttention
Main problemMoving large attention intermediates through GPU memoryWasted or duplicated storage for growing KV caches
Core ideaTile attention computation and reduce memory trafficMap logical token blocks to physical cache blocks
Useful outcomeMore efficient attention executionMore usable cache capacity for concurrent requests
What it does not doMake model weights smallerMake model weights smaller or remove attention computation
Can they work together?Yes, a compatible kernel can read a paged cacheYes, paging can be paired with efficient attention kernels

The original FlashAttention paper focuses on attention’s data movement. The PagedAttention paper focuses on dynamically managing KV state in a serving system. They are useful at different layers of the same request path.

What does FlashAttention optimize?

Standard attention computes softmax(Q × Kᵀ / √d) × V. A straightforward implementation writes a large matrix of scores, then probabilities, to GPU memory. FlashAttention tiles the calculation, uses fast on-chip storage, and accumulates the result without materializing the entire matrix in high-bandwidth memory. It computes exact attention in the algorithmic sense, with normal floating-point differences. FlashAttention research.

For one sequence with 32 query heads, an 8,192-token dense score tensor in a two-byte dtype occupies 4 GiB. At 16,384 tokens it occupies 16 GiB. Those are calculated sizes for one hypothetical full tensor, not measured peak usage: masking, fused kernels, intermediate precision, and the implementation change actual storage.

Avoiding that tensor does not make dense attention’s pairwise arithmetic linear. It reduces a major data-movement cost. It also does not eliminate the persistent KV state needed during generation. Follow the foundations course for the attention equations and the decode loop.

What does PagedAttention optimize?

An LLM server holds a growing KV cache for each active request. Reserving a large contiguous region for every possible output wastes capacity. PagedAttention divides the cache into fixed-size blocks and uses a block table to locate them, so a request can grow without requiring contiguous physical storage. Compatible blocks can also be shared. PagedAttention research.

Paging changes allocation, not the number of values required to represent an unshared token. There is still space overhead in a partially filled final block, plus metadata. More usable cache capacity can let a scheduler admit a larger batch; whether that improves user latency depends on the traffic and scheduler. See continuous batching explained.

Calculate attention intermediates and paged cache storage

This Python example keeps the two memory questions separate. First it estimates one dense score tensor. Then it compares reserving 128 tokens per request with allocating sixteen-token cache blocks for three invented request lengths. All numbers are illustrative; no GPU or model is loaded.

Run with Python 3; no GPU or packages required
# One full score tensor: batch 1, 32 query heads, 2 bytes per element.
for tokens in (4096, 8192, 16384):
    score_bytes = 32 * tokens * tokens * 2
    print(f"{tokens:,} tokens: dense scores {score_bytes / 2**30:.1f} GiB")

# Separate example: KV tensors for active requests, without block sharing.
lengths = [19, 45, 70]
block_size = 16
reserved_per_request = 128
kv_bytes_per_token = 128 * 1024  # Chosen teaching assumption, all layers.

reserved_tokens = len(lengths) * reserved_per_request
paged_tokens = sum(
    ((length + block_size - 1) // block_size) * block_size
    for length in lengths
)
used_tokens = sum(lengths)
for label, tokens in (("Fixed reservation", reserved_tokens),
                      ("Paged allocation", paged_tokens),
                      ("Used KV tensors", used_tokens)):
    mib = tokens * kv_bytes_per_token / 2**20
    print(f"{label}: {tokens} token slots, {mib:.2f} MiB")
print(f"Unused paged slots: {paged_tokens - used_tokens}")
Calculated output
4,096 tokens: dense scores 1.0 GiB
8,192 tokens: dense scores 4.0 GiB
16,384 tokens: dense scores 16.0 GiB
Fixed reservation: 384 token slots, 48.00 MiB
Paged allocation: 160 token slots, 20.00 MiB
Used KV tensors: 134 token slots, 16.75 MiB
Unused paged slots: 26

The paged version rounds 19 tokens to 32 slots, 45 to 48, and 70 to 80. It allocates 20 MiB rather than 48 MiB while keeping the same 134 tokens of state. That saves reserved capacity in this example; it is not a speed benchmark. The two parts measure different objects, so do not subtract the score-tensor estimate from the KV-cache budget.

To size an actual server, use its layer count, KV heads, head dimension, cache dtype, and active lengths. Query heads and KV heads can differ with grouped-query attention. GPU memory for LLM inference covers weights, KV state, and runtime headroom together.

Can FlashAttention and PagedAttention work together?

Yes. The FlashAttention project documents support for a paged KV cache through a block_table in its cache-aware interface. This is direct evidence that efficient attention execution and noncontiguous cache storage can be combined. Support still depends on the specific interface and backend. FlashAttention implementation.

An inference engine chooses an attention backend that works with the model, device, dtype, and attention features. vLLM publishes a backend support matrix covering constraints such as head sizes, cache dtypes, sliding windows, and attention variants. Check that matrix for your release before forcing a backend.

How prefill and decode change the diagnosis

Prefill processes many prompt positions, while a decode step usually adds one position per active request. That changes the shapes of the attention work. The FlashAttention repository includes a cache-aware inference interface for unequal query and cache lengths, so describing FlashAttention as only a training technique or only a prefill technique is misleading. FlashAttention cache interface.

Starting hypotheses to investigate, not automatic fixes.
Observed symptomWhat to inspectA useful next experiment
Long prompts start slowlyPrefill time, selected kernel, and queueingCompare prompt-length buckets at the same request rate
Few concurrent requests fitWeight footprint, KV budget, active lengths, and preemptionEstimate cache capacity and compare it with engine logs
Repeated long context is expensiveShared token prefixes and cache reuseCompare cold and warm runs with prefix caching
Answers stream slowlyDecode time, batch size, memory traffic, and communicationProfile the serving step before changing a kernel

Treat these as hypotheses because the same symptom can have several causes. LLM inference metrics explains how to separate time to first token from inter-token latency. Prefix caching explained covers the repeated-context experiment.

FlashAttention and PagedAttention FAQ

Which one should I enable first? Start from the serving engine’s supported defaults, record its backend, and establish a workload baseline. Use a measured bottleneck to decide what to change.

Are their published speedups directly comparable? No. Kernel timing, training time, and serving throughput describe different experiments. Match workload, hardware, quality, and latency constraints before drawing a conclusion.

Will they solve a model-weight out-of-memory error? Neither changes the weight tensor size. First budget the weights and runtime state, then consider quantization or model parallelism.

Where can I study the implementation? Use the vLLM course for cache allocation and scheduling, and the PyTorch course for tensors and attention operations.

Sources and further reading

Primary documentation and research behind this guide.

  1. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
  2. Efficient Memory Management for Large Language Model Serving with PagedAttention
  3. Dao-AILab: FlashAttention implementation and cache interface
  4. vLLM v0.30.0: attention backend feature support

Keep learning

Related guidePrefix caching explained: faster LLM prompts with KV reuseRelated guideGPU memory for LLM inference: how to size and choose a GPURelated guideContinuous batching explained: how LLM servers scaleRelated guideKV cache explained: the formula, a diagram, and a memory exampleAcademy courseInference Engineering FoundationsAcademy coursevLLMAcademy coursePyTorch
Back to all articles