Infrastructure6 min read

AI cloud infrastructure for LLM inference: a practical guide

Serving a large language model in the cloud takes more than a GPU and a container. This guide walks through each layer of AI inference infrastructure, from accelerators and serving engines to routing, autoscaling, and observability, and shows how to work out what each million tokens really costs.

The layers of an AI inference stack

AI cloud infrastructure for inference is the hardware and software that turns a trained model into a reliable API. Whether you build it yourself or rent it, the same layers are always present:

From the chip up to the dashboard. Each layer can become the bottleneck.
LayerWhat it doesExamples
AcceleratorsHold the weights and KV cache and run the modelNVIDIA L40S, H100, H200
Serving engineBatches requests, manages the KV cache, streams tokensvLLM, SGLang, TensorRT-LLM
OrchestrationPlaces engine replicas on GPU machines and restarts themKubernetes, managed GPU endpoints
RoutingSends each request to the replica that can serve it bestLoad balancer, cache-aware router
AutoscalingAdds or removes replicas as traffic changesKubernetes autoscalers, Dynamo planner
ObservabilityMeasures latency, throughput, queueing, and errorsPrometheus metrics, tracing, dashboards

Cloud deployment options for inference

More control means more work, and usually a lower cost per token at steady load.
OptionYou manageBest forWatch out for
Hosted model APIPrompts onlyPrototypes, spiky or low trafficPer-token prices at scale, limited model choice
Managed GPU endpointModel and engine settingsOpen models without running a clusterCold starts, less control over scheduling
Dedicated GPU instancesMachines, engine, and deploymentSteady traffic on one or a few modelsIdle GPUs you still pay for
Kubernetes GPU clusterThe whole stackMany models, teams, or regionsOperational complexity

A common path is to start on an API, move steady workloads to dedicated GPUs once traffic is predictable, and adopt a cluster when several models or teams share the hardware. The deciding numbers are your traffic shape and your cost per million tokens, covered at the end of this guide.

Choosing GPUs for inference

Pick the GPU from the model, not the other way round. Memory capacity decides whether the weights fit and how many requests share the GPU; memory bandwidth decides how quickly each answer streams. An 8B model fits comfortably on a 48 GB L40S, while a 32B model in BF16 needs an 80 GB GPU or more to serve more than a handful of users. GPU memory for LLM inference includes a calculator for these estimates.

Benchmark on the exact instance type before committing. Two clouds can attach the same GPU to different CPUs, networking, and local storage, and those differences show up in model load time and multi-GPU performance.

Scaling from one GPU to many

There are two separate reasons to add GPUs, and they call for different techniques:

  • The model does not fit. Split it with tensor parallelism across the GPUs of one machine, then add pipeline parallelism across machines only if one node is not enough. The vLLM guide recommends exactly this order. vLLM: parallelism and scaling.
  • The model fits but traffic does not. Run more independent replicas behind a router. Each replica serves its own requests, so throughput grows roughly in proportion to the number of replicas.

Prefer replicas whenever the model fits on one GPU. Tensor parallelism adds communication between GPUs at every layer, so two GPUs running one split model usually deliver less than twice the throughput of one.

Running inference on Kubernetes

Kubernetes schedules GPUs as an extended resource exposed by a device plugin. A container requests whole GPUs in its resource limits. Kubernetes: schedule GPUs.

A minimal vLLM pod that requests one GPU
apiVersion: v1
kind: Pod
metadata:
  name: vllm-qwen3-8b
spec:
  containers:
    - name: vllm
      image: vllm/vllm-openai:v0.30.0
      args: ["--model", "Qwen/Qwen3-8B", "--max-model-len", "8192"]
      ports:
        - containerPort: 8000
      readinessProbe:
        httpGet:
          path: /health
          port: 8000
      resources:
        limits:
          nvidia.com/gpu: 1

A production setup adds a Deployment for restarts, a persistent volume or model cache so each new pod does not download the weights again, and an authenticated service in front. The vLLM Kubernetes guide covers these pieces. The readiness probe matters: loading weights can take minutes, and traffic should not reach a pod until it answers on /health.

Routing and autoscaling for LLM traffic

Round-robin load balancing ignores what LLM servers care about. Requests that share a long system prompt or document run faster on the replica that already holds that prefix in its KV cache. Cache-aware routers send them there. NVIDIA’s open-source Dynamo routes on KV cache overlap and worker load across vLLM, SGLang, and TensorRT-LLM workers.

Scale on signals that reflect the GPU’s real state, not CPU usage. vLLM exports vllm:num_requests_waiting, vllm:num_requests_running, and vllm:kv_cache_usage_perc, along with time to first token and inter-token latency histograms. vLLM: production metrics. A growing waiting queue or a full KV cache is a better reason to add a replica than a busy CPU.

Disaggregated prefill and decode

Prefill is compute-heavy and decode is memory-bandwidth-heavy. Disaggregated serving runs them on separate GPU pools and moves the KV cache between them, so each pool can be sized and tuned for its phase. DistServe and Splitwise introduced the approach, and Dynamo builds on it.

It is a latency tool, not a throughput shortcut. The vLLM documentation states that disaggregated prefill does not improve throughput; its benefit is tuning time to first token and inter-token latency independently and controlling tail latency. vLLM: disaggregated prefilling. Start with one engine per replica and consider it once latency targets are hard to meet.

Calculate cost per million tokens

Self-hosted cost is the GPU price divided by the tokens you actually serve. The script below uses an illustrative $3.00 per GPU-hour and illustrative throughput figures; replace both with your provider’s price and your own benchmark. LLM inference metrics explains how to measure throughput correctly.

Run with Python 3; no packages required
# Cost per million output tokens from a GPU price and a measured throughput.
# Both inputs below are illustrative; use your provider's price and your benchmark.


def cost_per_million(gpu_hourly_usd, gpus, tokens_per_second, utilization):
    """utilization: the share of each hour the replica does useful work."""
    hourly = gpu_hourly_usd * gpus
    tokens_per_hour = tokens_per_second * 3600 * utilization
    return hourly / tokens_per_hour * 1e6


scenarios = [
    ("1 GPU, quiet traffic", 1, 2500, 0.15),
    ("1 GPU, busy", 1, 2500, 0.70),
    ("2 GPUs, tensor parallel", 2, 4000, 0.70),
]

for name, gpus, tps, util in scenarios:
    usd = cost_per_million(gpu_hourly_usd=3.00, gpus=gpus, tokens_per_second=tps, utilization=util)
    print(f"{name:<24} ${usd:5.2f} per million output tokens")
Calculated output
1 GPU, quiet traffic     $ 2.22 per million output tokens
1 GPU, busy              $ 0.48 per million output tokens
2 GPUs, tensor parallel  $ 0.60 per million output tokens

Utilization dominates. The same GPU costs more than four times as much per token when it sits mostly idle, which is why steady traffic favors dedicated GPUs and spiky traffic favors APIs or scale-to-zero endpoints. The tensor parallel row shows the other lesson from earlier: splitting a model that already fits across two GPUs raised throughput but also raised the cost per token. Compare the result with the per-token price of a hosted API for the same model before deciding to self-host.

AI inference infrastructure FAQ

Is self-hosting cheaper than an API? Only when your GPUs stay busy. At low or unpredictable traffic, an API’s per-token price is often lower than paying for idle hardware.

Do I need Kubernetes to serve LLMs? No. One model on one machine runs fine under Docker or a managed endpoint. Kubernetes pays off when you operate several models, replicas, or teams on shared GPUs.

Which serving engine should I use? vLLM and SGLang are the most widely used open-source engines. vLLM vs SGLang compares them, and the vLLM tutorial takes you from weights to a working endpoint.

How do I learn to run this in production? The vLLM and SGLang courses cover single-engine serving, and the NVIDIA Dynamo course covers routing and disaggregated serving across a fleet. How to learn inference engineering puts them in order.

Sources and further reading

Primary documentation and research behind this guide.

  1. vLLM: parallelism and scaling
  2. Kubernetes: schedule GPUs
  3. vLLM: deploying with Kubernetes
  4. NVIDIA Dynamo
  5. vLLM: production metrics
  6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving (2024)
  7. Splitwise: Efficient Generative LLM Inference Using Phase Splitting (2023)
  8. vLLM: disaggregated prefilling

Keep learning

Related guideGPU memory for LLM inference: how to size and choose a GPURelated guideWhat is AI inference? How trained models answer in productionRelated guideContinuous batching explained: how LLM servers scaleRelated guideLLM inference metrics: TTFT, TPOT, ITL, and throughputAcademy coursevLLMAcademy courseSGLangAcademy courseNVIDIA Dynamo
Back to all articles