What is AI inference?
AI inference is running a trained model on new input to produce an output: a label, a score, an image, or the next words of an answer. The model’s weights were learned earlier, during training. At inference time they stay fixed; the model only reads them to turn input into a prediction.
A spam filter deciding whether an email is junk, a vision model finding a tumor in a scan, and a chat assistant writing a reply are all doing inference. In each case the expensive learning already happened. What remains is answering requests quickly, correctly, and at a cost the product can afford, often millions of times a day.
Training vs inference
Training and inference use the same model but are very different workloads. Training adjusts the weights by showing the model data and correcting its mistakes. Inference uses the finished weights to answer real requests.
| Aspect | Training | Inference |
|---|---|---|
| Goal | Learn weights from data | Use fixed weights to answer requests |
| How often it runs | Once per model version, for days or weeks | Continuously, for every request, for as long as the model is in use |
| Weights | Updated after each batch | Read only |
| What success looks like | Lower loss and better evaluation scores | Correct answers within a latency and cost target |
| Main constraint | Total compute and time to finish | Latency per request, throughput, and memory |
The second row explains why inference matters so much commercially. A model is trained a limited number of times, but it serves requests for its whole lifetime. Small savings per request, multiplied across all traffic, decide whether a product is profitable.
Training and inference in a few lines of Python
This toy “language model” learns which word tends to follow which, then uses those fixed counts to continue a prompt. Real models learn billions of weights instead of a table of counts, but the split is the same: learning happens once, and inference reads what was learned.
from collections import Counter, defaultdict
# Training: count which word follows which. This happens once, offline.
corpus = "the gpu reads the weights . the gpu writes the answer . the server streams the answer ."
words = corpus.split()
follows = defaultdict(Counter)
for current, nxt in zip(words, words[1:]):
follows[current][nxt] += 1
# Inference: the counts are now fixed; we only read them to predict.
def generate(prompt, max_new_tokens=6):
tokens = prompt.split()
for _ in range(max_new_tokens):
options = follows.get(tokens[-1])
if not options:
break
next_token = options.most_common(1)[0][0] # greedy: pick the top score
tokens.append(next_token)
if next_token == ".":
break
return " ".join(tokens)
print("learned options after 'the':", dict(follows["the"]))
print(generate("the server"))
print(generate("gpu"))learned options after 'the': {'gpu': 2, 'weights': 1, 'answer': 2, 'server': 1}
the server streams the gpu reads the gpu
gpu reads the gpu reads the gpuTwo lessons from real inference already show up here. First, generation is a loop: the model predicts one token, appends it, and predicts again. Second, the decoding rule matters. Always choosing the top option makes this tiny model repeat itself until it hits the max_new_tokens limit. Production servers expose sampling settings and stop conditions for the same reason. Hugging Face’s generation guide describes these controls.
How LLM inference works
Large language models are built on the transformer architecture. When a request arrives, the server turns the text into tokens and runs the model in two phases:
- Prefill processes the whole prompt at once. It is compute-heavy and determines the time to first token.
- Decode generates the answer one token at a time. Each step reads all of the model’s weights from GPU memory, so it is usually limited by memory bandwidth rather than arithmetic.
To avoid recomputing the whole conversation at every step, the model stores intermediate attention results in a KV cache. That cache grows with every token and every concurrent user, and it is often what limits how many requests a GPU can serve. LLM inference explained follows a request through these stages, and the KV cache explained shows how to calculate its size.
What makes inference fast or expensive
Three resources set the speed and price of LLM inference:
- GPU memory capacity decides whether the weights fit and how much room is left for the KV cache. More cache room means more users at once.
- Memory bandwidth caps how fast a single answer can stream, because every decode step reads the weights.
- Utilization decides cost. A GPU billed by the hour costs the same whether it serves one request or a hundred, so servers batch many requests together.
Engines such as vLLM manage the cache in pages to fit more requests in memory, an idea introduced in the PagedAttention paper, and schedule requests with continuous batching. To judge whether those choices help, measure time to first token, time per output token, and throughput together. LLM inference metrics defines each one.
Where AI inference runs
| Option | What you control | Typical fit |
|---|---|---|
| Hosted model API | Prompts and parameters only | Fast start, variable or low traffic |
| Cloud GPUs you rent | Model, engine, and scaling | Steady traffic, custom or open models, data control |
| Your own data center | Everything, including hardware | Very large, predictable workloads |
| On-device or edge | A small model on a phone, laptop, or embedded chip | Offline use, privacy, very low latency |
Most teams that serve open models use cloud GPUs. AI cloud infrastructure for LLM inference explains how that stack fits together, and GPU memory for LLM inference shows how to pick a GPU for a given model.
How inference is optimized
Inference optimization means producing the same quality of answer with less time, memory, or money. The most widely used techniques are:
- Quantization: storing weights in fewer bits so they take less memory and are faster to read.
- Speculative decoding: letting a cheap draft propose several tokens that the main model checks in one step.
- Continuous batching: adding and removing requests from the running batch at every step.
- Prefix caching: reusing the KV cache of a shared system prompt or document across requests. See vLLM vs SGLang for how two engines approach it.
Each of these trades something: quantization can affect accuracy, and larger batches can slow individual streams. Change one at a time, measure, and check answer quality alongside speed. The vLLM optimization guide lists the settings that control these tradeoffs in one engine.
AI inference FAQ
Is inference cheaper than training? Per request, yes, by a wide margin. In total it can cost more, because a popular model answers requests every second for as long as it is deployed.
Does a model learn during inference? No. The weights are fixed. A chat model “remembers” earlier messages only because they are sent again as part of the input.
Do I need a GPU for inference? Not always. Small models run on CPUs and phones. Large language models are usually served on GPUs or other accelerators because they need high memory bandwidth.
What is an inference engineer? An engineer who makes trained models fast, reliable, and affordable to serve. How to learn inference engineering lays out the skills step by step, and the free Foundations course is a good place to start.
Sources and further reading
Primary documentation and research behind this guide.