The layers of an AI inference stack
AI cloud infrastructure for inference is the hardware and software that turns a trained model into a reliable API. Whether you build it yourself or rent it, the same layers are always present:
| Layer | What it does | Examples |
|---|---|---|
| Accelerators | Hold the weights and KV cache and run the model | NVIDIA L40S, H100, H200 |
| Serving engine | Batches requests, manages the KV cache, streams tokens | vLLM, SGLang, TensorRT-LLM |
| Orchestration | Places engine replicas on GPU machines and restarts them | Kubernetes, managed GPU endpoints |
| Routing | Sends each request to the replica that can serve it best | Load balancer, cache-aware router |
| Autoscaling | Adds or removes replicas as traffic changes | Kubernetes autoscalers, Dynamo planner |
| Observability | Measures latency, throughput, queueing, and errors | Prometheus metrics, tracing, dashboards |
Cloud deployment options for inference
| Option | You manage | Best for | Watch out for |
|---|---|---|---|
| Hosted model API | Prompts only | Prototypes, spiky or low traffic | Per-token prices at scale, limited model choice |
| Managed GPU endpoint | Model and engine settings | Open models without running a cluster | Cold starts, less control over scheduling |
| Dedicated GPU instances | Machines, engine, and deployment | Steady traffic on one or a few models | Idle GPUs you still pay for |
| Kubernetes GPU cluster | The whole stack | Many models, teams, or regions | Operational complexity |
A common path is to start on an API, move steady workloads to dedicated GPUs once traffic is predictable, and adopt a cluster when several models or teams share the hardware. The deciding numbers are your traffic shape and your cost per million tokens, covered at the end of this guide.
Choosing GPUs for inference
Pick the GPU from the model, not the other way round. Memory capacity decides whether the weights fit and how many requests share the GPU; memory bandwidth decides how quickly each answer streams. An 8B model fits comfortably on a 48 GB L40S, while a 32B model in BF16 needs an 80 GB GPU or more to serve more than a handful of users. GPU memory for LLM inference includes a calculator for these estimates.
Benchmark on the exact instance type before committing. Two clouds can attach the same GPU to different CPUs, networking, and local storage, and those differences show up in model load time and multi-GPU performance.
Scaling from one GPU to many
There are two separate reasons to add GPUs, and they call for different techniques:
- The model does not fit. Split it with tensor parallelism across the GPUs of one machine, then add pipeline parallelism across machines only if one node is not enough. The vLLM guide recommends exactly this order. vLLM: parallelism and scaling.
- The model fits but traffic does not. Run more independent replicas behind a router. Each replica serves its own requests, so throughput grows roughly in proportion to the number of replicas.
Prefer replicas whenever the model fits on one GPU. Tensor parallelism adds communication between GPUs at every layer, so two GPUs running one split model usually deliver less than twice the throughput of one.
Running inference on Kubernetes
Kubernetes schedules GPUs as an extended resource exposed by a device plugin. A container requests whole GPUs in its resource limits. Kubernetes: schedule GPUs.
apiVersion: v1
kind: Pod
metadata:
name: vllm-qwen3-8b
spec:
containers:
- name: vllm
image: vllm/vllm-openai:v0.30.0
args: ["--model", "Qwen/Qwen3-8B", "--max-model-len", "8192"]
ports:
- containerPort: 8000
readinessProbe:
httpGet:
path: /health
port: 8000
resources:
limits:
nvidia.com/gpu: 1A production setup adds a Deployment for restarts, a persistent volume or model cache so each new pod does not download the weights again, and an authenticated service in front. The vLLM Kubernetes guide covers these pieces. The readiness probe matters: loading weights can take minutes, and traffic should not reach a pod until it answers on /health.
Routing and autoscaling for LLM traffic
Round-robin load balancing ignores what LLM servers care about. Requests that share a long system prompt or document run faster on the replica that already holds that prefix in its KV cache. Cache-aware routers send them there. NVIDIA’s open-source Dynamo routes on KV cache overlap and worker load across vLLM, SGLang, and TensorRT-LLM workers.
Scale on signals that reflect the GPU’s real state, not CPU usage. vLLM exports vllm:num_requests_waiting, vllm:num_requests_running, and vllm:kv_cache_usage_perc, along with time to first token and inter-token latency histograms. vLLM: production metrics. A growing waiting queue or a full KV cache is a better reason to add a replica than a busy CPU.
Disaggregated prefill and decode
Prefill is compute-heavy and decode is memory-bandwidth-heavy. Disaggregated serving runs them on separate GPU pools and moves the KV cache between them, so each pool can be sized and tuned for its phase. DistServe and Splitwise introduced the approach, and Dynamo builds on it.
It is a latency tool, not a throughput shortcut. The vLLM documentation states that disaggregated prefill does not improve throughput; its benefit is tuning time to first token and inter-token latency independently and controlling tail latency. vLLM: disaggregated prefilling. Start with one engine per replica and consider it once latency targets are hard to meet.
Calculate cost per million tokens
Self-hosted cost is the GPU price divided by the tokens you actually serve. The script below uses an illustrative $3.00 per GPU-hour and illustrative throughput figures; replace both with your provider’s price and your own benchmark. LLM inference metrics explains how to measure throughput correctly.
# Cost per million output tokens from a GPU price and a measured throughput.
# Both inputs below are illustrative; use your provider's price and your benchmark.
def cost_per_million(gpu_hourly_usd, gpus, tokens_per_second, utilization):
"""utilization: the share of each hour the replica does useful work."""
hourly = gpu_hourly_usd * gpus
tokens_per_hour = tokens_per_second * 3600 * utilization
return hourly / tokens_per_hour * 1e6
scenarios = [
("1 GPU, quiet traffic", 1, 2500, 0.15),
("1 GPU, busy", 1, 2500, 0.70),
("2 GPUs, tensor parallel", 2, 4000, 0.70),
]
for name, gpus, tps, util in scenarios:
usd = cost_per_million(gpu_hourly_usd=3.00, gpus=gpus, tokens_per_second=tps, utilization=util)
print(f"{name:<24} ${usd:5.2f} per million output tokens")1 GPU, quiet traffic $ 2.22 per million output tokens
1 GPU, busy $ 0.48 per million output tokens
2 GPUs, tensor parallel $ 0.60 per million output tokensUtilization dominates. The same GPU costs more than four times as much per token when it sits mostly idle, which is why steady traffic favors dedicated GPUs and spiky traffic favors APIs or scale-to-zero endpoints. The tensor parallel row shows the other lesson from earlier: splitting a model that already fits across two GPUs raised throughput but also raised the cost per token. Compare the result with the per-token price of a hosted API for the same model before deciding to self-host.
AI inference infrastructure FAQ
Is self-hosting cheaper than an API? Only when your GPUs stay busy. At low or unpredictable traffic, an API’s per-token price is often lower than paying for idle hardware.
Do I need Kubernetes to serve LLMs? No. One model on one machine runs fine under Docker or a managed endpoint. Kubernetes pays off when you operate several models, replicas, or teams on shared GPUs.
Which serving engine should I use? vLLM and SGLang are the most widely used open-source engines. vLLM vs SGLang compares them, and the vLLM tutorial takes you from weights to a working endpoint.
How do I learn to run this in production? The vLLM and SGLang courses cover single-engine serving, and the NVIDIA Dynamo course covers routing and disaggregated serving across a fleet. How to learn inference engineering puts them in order.
Sources and further reading
Primary documentation and research behind this guide.
- vLLM: parallelism and scaling
- Kubernetes: schedule GPUs
- vLLM: deploying with Kubernetes
- NVIDIA Dynamo
- vLLM: production metrics
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving (2024)
- Splitwise: Efficient Generative LLM Inference Using Phase Splitting (2023)
- vLLM: disaggregated prefilling