What inference engineers do
An inference engineer owns the path from a trained model to answers in production. Day to day that means choosing hardware, configuring serving engines such as vLLM and SGLang, measuring latency and throughput, applying optimizations like quantization and speculative decoding, and keeping the service reliable as traffic grows. The goal is always the same: the right answer, quickly, at a cost the product can sustain.
The role exists because inference is where most of a deployed model’s cost and user experience live. If you are new to the idea, read what is AI inference first, then come back to this roadmap.
Prerequisites
You do not need a research background, but you need working comfort with:
- Python, including reading other people’s code and writing small scripts to measure things.
- Linux, Docker, and HTTP APIs, because every serving engine runs as a containerized service behind an API.
- Tensors and models at a basic level: what a forward pass is and how weights are stored. The PyTorch course and the Hugging Face Transformers course cover this.
- A rough mental model of a GPU: it has its own memory, that memory has a capacity and a bandwidth, and moving data is often slower than computing on it.
The inference engineering roadmap
| Stage | What to learn | Where to start |
|---|---|---|
| 1. Fundamentals | Tokens, prefill and decode, the KV cache, and serving metrics | Foundations course, LLM inference explained |
| 2. Serve a model | Run an engine, call its API, and benchmark it | vLLM tutorial, vLLM course, SGLang course |
| 3. Optimize | Batching, quantization, speculative decoding, prefix caching | Continuous batching, quantization, speculative decoding |
| 4. Scale | GPU sizing, parallelism, routing, autoscaling, cost | GPU memory guide, cloud infrastructure guide, Dynamo course |
| 5. Go deeper | Kernels, attention implementations, and profiling | Papers and references on the resources page |
Stage 1: learn how inference works
Start with the mechanics, because every later decision depends on them. Learn how a prompt becomes tokens, why generation splits into a compute-heavy prefill and a memory-bound decode, how the KV cache grows, and what TTFT, TPOT, and throughput measure.
The most useful habit to build early is the back-of-the-envelope estimate. During decode, each step reads every weight from GPU memory, so memory bandwidth divided by weight size gives an upper bound on tokens per second for a single request. This first exercise uses Qwen3-8B, which has about 8.2 billion parameters, and the bandwidth figures from NVIDIA’s L40S, H100, and H200 pages.
# Decode ceiling: each step reads every weight from GPU memory once.
GB = 1e9
params = 8.19e9 # Qwen3-8B
gpus = {"L40S": 864, "H100 SXM": 3350, "H200": 4800} # GB/s, from NVIDIA spec pages
for label, bytes_per_weight in (("BF16", 2), ("FP8", 1)):
weight_gb = params * bytes_per_weight / GB
print(f"{label} weights: {weight_gb:.1f} GB")
for gpu, bandwidth in gpus.items():
steps_per_second = bandwidth / weight_gb
print(f" {gpu:<9} at most {steps_per_second:4.0f} tokens/s for one request")BF16 weights: 16.4 GB
L40S at most 53 tokens/s for one request
H100 SXM at most 205 tokens/s for one request
H200 at most 293 tokens/s for one request
FP8 weights: 8.2 GB
L40S at most 105 tokens/s for one request
H100 SXM at most 409 tokens/s for one request
H200 at most 586 tokens/s for one requestThese are ceilings, not predictions: real engines also read the KV cache and never reach peak bandwidth. Still, the estimate explains two facts you will rely on constantly. Halving the bytes per weight roughly doubles single-stream speed, which is why quantization helps decode. And because one request leaves most of the GPU’s arithmetic idle, batching many requests together raises total throughput far more than it slows each one.
Stage 2: serve and benchmark a real model
Rent a single GPU, start an engine, and send it requests. Follow the vLLM tutorial to get a pinned, reproducible endpoint, then run a benchmark at several request rates and record TTFT, TPOT, and throughput at each. vLLM’s benchmark command is a good starting point.
Then repeat the exercise with SGLang. Running the same workload on two engines teaches you which results come from the model and which come from the scheduler. vLLM vs SGLang explains how to make that comparison fair.
Stage 3: optimize one change at a time
With a baseline in hand, apply one optimization at a time and measure both speed and answer quality:
- Tune continuous batching limits and watch the latency and throughput tradeoff.
- Serve an FP8 or 4-bit version of the model and compare accuracy on your own prompts, as described in LLM quantization explained.
- Enable speculative decoding and check its acceptance rate at low and high load.
- Turn on prefix caching for a workload with a long shared system prompt and measure the change in time to first token.
Stage 4: scale to production infrastructure
Production inference adds problems a single GPU never shows: models that do not fit, traffic that does not fit, cold starts, routing, and cost. Learn to size GPUs with GPU memory for LLM inference, then study tensor and pipeline parallelism, replicas, cache-aware routing, and autoscaling in AI cloud infrastructure for LLM inference. The NVIDIA Dynamo course puts these together for a fleet of engines.
Stage 5: go deeper into kernels and research
Once the system-level picture is clear, read the papers behind the tools you use. PagedAttention explains vLLM’s memory manager, and FlashAttention shows how an attention kernel is designed around GPU memory. Learn to read a profiler trace so you can see where a step spends its time. The resources page collects books, papers, and references for this stage.
Projects that prove inference skills
| Project | What it demonstrates |
|---|---|
| Benchmark one model on two engines at three request rates | Measurement discipline and knowledge of the metrics |
| Serve a quantized model and report speed and accuracy against BF16 | Optimization with quality checks |
| Find the request rate where TTFT exceeds a target on one GPU | Capacity planning |
| Calculate cost per million tokens for your setup and compare it with an API | Infrastructure and cost reasoning |
| Deploy two replicas behind a router with health checks | Production serving |
The coding practice problems give you smaller exercises along the way, and the jobs board shows what companies hiring inference engineers ask for.
Learning inference engineering FAQ
Do I need a machine learning degree? No. Inference engineering draws more on systems skills, such as measurement, memory, scheduling, and operations, than on model training. You need to understand how a model runs, not how to invent one.
Do I need my own GPU? No. Stage 1 runs on any laptop. For later stages, renting a cloud GPU by the hour is enough; shut it down between sessions.
How long does it take? It depends on your starting point. Developers who already know Python and Docker can work through the fundamentals and serve their first benchmarked model in a few weeks of steady practice; production-scale skills take longer and come from projects.
Where should I start? With the Foundations course, which is free, and LLM inference explained. See pricing for the full courses.
Sources and further reading
Primary documentation and research behind this guide.