Careers5 min read

How to learn inference engineering: a step-by-step roadmap

Inference engineering is the work of making trained models fast, reliable, and affordable to serve. It sits between machine learning and systems engineering, and few people learn it in school. This roadmap lays out what to learn, in what order, and what to build at each stage to prove you can do the job.

What inference engineers do

An inference engineer owns the path from a trained model to answers in production. Day to day that means choosing hardware, configuring serving engines such as vLLM and SGLang, measuring latency and throughput, applying optimizations like quantization and speculative decoding, and keeping the service reliable as traffic grows. The goal is always the same: the right answer, quickly, at a cost the product can sustain.

The role exists because inference is where most of a deployed model’s cost and user experience live. If you are new to the idea, read what is AI inference first, then come back to this roadmap.

Prerequisites

You do not need a research background, but you need working comfort with:

  • Python, including reading other people’s code and writing small scripts to measure things.
  • Linux, Docker, and HTTP APIs, because every serving engine runs as a containerized service behind an API.
  • Tensors and models at a basic level: what a forward pass is and how weights are stored. The PyTorch course and the Hugging Face Transformers course cover this.
  • A rough mental model of a GPU: it has its own memory, that memory has a capacity and a bandwidth, and moving data is often slower than computing on it.

The inference engineering roadmap

Learn each stage before the next. Every stage ends with something you can run and measure.
StageWhat to learnWhere to start
1. FundamentalsTokens, prefill and decode, the KV cache, and serving metricsFoundations course, LLM inference explained
2. Serve a modelRun an engine, call its API, and benchmark itvLLM tutorial, vLLM course, SGLang course
3. OptimizeBatching, quantization, speculative decoding, prefix cachingContinuous batching, quantization, speculative decoding
4. ScaleGPU sizing, parallelism, routing, autoscaling, costGPU memory guide, cloud infrastructure guide, Dynamo course
5. Go deeperKernels, attention implementations, and profilingPapers and references on the resources page

Stage 1: learn how inference works

Start with the mechanics, because every later decision depends on them. Learn how a prompt becomes tokens, why generation splits into a compute-heavy prefill and a memory-bound decode, how the KV cache grows, and what TTFT, TPOT, and throughput measure.

The most useful habit to build early is the back-of-the-envelope estimate. During decode, each step reads every weight from GPU memory, so memory bandwidth divided by weight size gives an upper bound on tokens per second for a single request. This first exercise uses Qwen3-8B, which has about 8.2 billion parameters, and the bandwidth figures from NVIDIA’s L40S, H100, and H200 pages.

Run with Python 3; no GPU or packages required
# Decode ceiling: each step reads every weight from GPU memory once.
GB = 1e9

params = 8.19e9  # Qwen3-8B
gpus = {"L40S": 864, "H100 SXM": 3350, "H200": 4800}  # GB/s, from NVIDIA spec pages

for label, bytes_per_weight in (("BF16", 2), ("FP8", 1)):
    weight_gb = params * bytes_per_weight / GB
    print(f"{label} weights: {weight_gb:.1f} GB")
    for gpu, bandwidth in gpus.items():
        steps_per_second = bandwidth / weight_gb
        print(f"  {gpu:<9} at most {steps_per_second:4.0f} tokens/s for one request")
Calculated output
BF16 weights: 16.4 GB
  L40S      at most   53 tokens/s for one request
  H100 SXM  at most  205 tokens/s for one request
  H200      at most  293 tokens/s for one request
FP8 weights: 8.2 GB
  L40S      at most  105 tokens/s for one request
  H100 SXM  at most  409 tokens/s for one request
  H200      at most  586 tokens/s for one request

These are ceilings, not predictions: real engines also read the KV cache and never reach peak bandwidth. Still, the estimate explains two facts you will rely on constantly. Halving the bytes per weight roughly doubles single-stream speed, which is why quantization helps decode. And because one request leaves most of the GPU’s arithmetic idle, batching many requests together raises total throughput far more than it slows each one.

Stage 2: serve and benchmark a real model

Rent a single GPU, start an engine, and send it requests. Follow the vLLM tutorial to get a pinned, reproducible endpoint, then run a benchmark at several request rates and record TTFT, TPOT, and throughput at each. vLLM’s benchmark command is a good starting point.

Then repeat the exercise with SGLang. Running the same workload on two engines teaches you which results come from the model and which come from the scheduler. vLLM vs SGLang explains how to make that comparison fair.

Stage 3: optimize one change at a time

With a baseline in hand, apply one optimization at a time and measure both speed and answer quality:

  • Tune continuous batching limits and watch the latency and throughput tradeoff.
  • Serve an FP8 or 4-bit version of the model and compare accuracy on your own prompts, as described in LLM quantization explained.
  • Enable speculative decoding and check its acceptance rate at low and high load.
  • Turn on prefix caching for a workload with a long shared system prompt and measure the change in time to first token.

Stage 4: scale to production infrastructure

Production inference adds problems a single GPU never shows: models that do not fit, traffic that does not fit, cold starts, routing, and cost. Learn to size GPUs with GPU memory for LLM inference, then study tensor and pipeline parallelism, replicas, cache-aware routing, and autoscaling in AI cloud infrastructure for LLM inference. The NVIDIA Dynamo course puts these together for a fleet of engines.

Stage 5: go deeper into kernels and research

Once the system-level picture is clear, read the papers behind the tools you use. PagedAttention explains vLLM’s memory manager, and FlashAttention shows how an attention kernel is designed around GPU memory. Learn to read a profiler trace so you can see where a step spends its time. The resources page collects books, papers, and references for this stage.

Projects that prove inference skills

Each project produces numbers you can explain in an interview.
ProjectWhat it demonstrates
Benchmark one model on two engines at three request ratesMeasurement discipline and knowledge of the metrics
Serve a quantized model and report speed and accuracy against BF16Optimization with quality checks
Find the request rate where TTFT exceeds a target on one GPUCapacity planning
Calculate cost per million tokens for your setup and compare it with an APIInfrastructure and cost reasoning
Deploy two replicas behind a router with health checksProduction serving

The coding practice problems give you smaller exercises along the way, and the jobs board shows what companies hiring inference engineers ask for.

Learning inference engineering FAQ

Do I need a machine learning degree? No. Inference engineering draws more on systems skills, such as measurement, memory, scheduling, and operations, than on model training. You need to understand how a model runs, not how to invent one.

Do I need my own GPU? No. Stage 1 runs on any laptop. For later stages, renting a cloud GPU by the hour is enough; shut it down between sessions.

How long does it take? It depends on your starting point. Developers who already know Python and Docker can work through the fundamentals and serve their first benchmarked model in a few weeks of steady practice; production-scale skills take longer and come from projects.

Where should I start? With the Foundations course, which is free, and LLM inference explained. See pricing for the full courses.

Sources and further reading

Primary documentation and research behind this guide.

  1. NVIDIA L40S specifications
  2. NVIDIA H100 specifications
  3. NVIDIA H200 specifications
  4. vLLM: bench serve
  5. Efficient Memory Management with PagedAttention (2023)
  6. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (2022)

Keep learning

Related guideLLM inference explained: from prompt to generated tokensRelated guideGPU memory for LLM inference: how to size and choose a GPURelated guideAI cloud infrastructure for LLM inference: a practical guideRelated guideWhat is AI inference? How trained models answer in productionAcademy courseInference Engineering FoundationsAcademy coursevLLMAcademy courseSGLangAcademy courseNVIDIA Dynamo
Back to all articles