Foundations5 min read

What is AI inference? How trained models answer in production

Training gets the headlines, but every chatbot reply, search ranking, and code suggestion is produced by inference. This guide explains what AI inference is, how it differs from training, how large language models run it, and why it has become an engineering discipline of its own.

What is AI inference?

AI inference is running a trained model on new input to produce an output: a label, a score, an image, or the next words of an answer. The model’s weights were learned earlier, during training. At inference time they stay fixed; the model only reads them to turn input into a prediction.

A spam filter deciding whether an email is junk, a vision model finding a tumor in a scan, and a chat assistant writing a reply are all doing inference. In each case the expensive learning already happened. What remains is answering requests quickly, correctly, and at a cost the product can afford, often millions of times a day.

Training vs inference

Training and inference use the same model but are very different workloads. Training adjusts the weights by showing the model data and correcting its mistakes. Inference uses the finished weights to answer real requests.

How the two phases of a model’s life differ.
AspectTrainingInference
GoalLearn weights from dataUse fixed weights to answer requests
How often it runsOnce per model version, for days or weeksContinuously, for every request, for as long as the model is in use
WeightsUpdated after each batchRead only
What success looks likeLower loss and better evaluation scoresCorrect answers within a latency and cost target
Main constraintTotal compute and time to finishLatency per request, throughput, and memory

The second row explains why inference matters so much commercially. A model is trained a limited number of times, but it serves requests for its whole lifetime. Small savings per request, multiplied across all traffic, decide whether a product is profitable.

Training and inference in a few lines of Python

This toy “language model” learns which word tends to follow which, then uses those fixed counts to continue a prompt. Real models learn billions of weights instead of a table of counts, but the split is the same: learning happens once, and inference reads what was learned.

Run with Python 3; no packages required
from collections import Counter, defaultdict

# Training: count which word follows which. This happens once, offline.
corpus = "the gpu reads the weights . the gpu writes the answer . the server streams the answer ."
words = corpus.split()
follows = defaultdict(Counter)
for current, nxt in zip(words, words[1:]):
    follows[current][nxt] += 1


# Inference: the counts are now fixed; we only read them to predict.
def generate(prompt, max_new_tokens=6):
    tokens = prompt.split()
    for _ in range(max_new_tokens):
        options = follows.get(tokens[-1])
        if not options:
            break
        next_token = options.most_common(1)[0][0]  # greedy: pick the top score
        tokens.append(next_token)
        if next_token == ".":
            break
    return " ".join(tokens)


print("learned options after 'the':", dict(follows["the"]))
print(generate("the server"))
print(generate("gpu"))
Output
learned options after 'the': {'gpu': 2, 'weights': 1, 'answer': 2, 'server': 1}
the server streams the gpu reads the gpu
gpu reads the gpu reads the gpu

Two lessons from real inference already show up here. First, generation is a loop: the model predicts one token, appends it, and predicts again. Second, the decoding rule matters. Always choosing the top option makes this tiny model repeat itself until it hits the max_new_tokens limit. Production servers expose sampling settings and stop conditions for the same reason. Hugging Face’s generation guide describes these controls.

How LLM inference works

Large language models are built on the transformer architecture. When a request arrives, the server turns the text into tokens and runs the model in two phases:

  • Prefill processes the whole prompt at once. It is compute-heavy and determines the time to first token.
  • Decode generates the answer one token at a time. Each step reads all of the model’s weights from GPU memory, so it is usually limited by memory bandwidth rather than arithmetic.

To avoid recomputing the whole conversation at every step, the model stores intermediate attention results in a KV cache. That cache grows with every token and every concurrent user, and it is often what limits how many requests a GPU can serve. LLM inference explained follows a request through these stages, and the KV cache explained shows how to calculate its size.

What makes inference fast or expensive

Three resources set the speed and price of LLM inference:

  • GPU memory capacity decides whether the weights fit and how much room is left for the KV cache. More cache room means more users at once.
  • Memory bandwidth caps how fast a single answer can stream, because every decode step reads the weights.
  • Utilization decides cost. A GPU billed by the hour costs the same whether it serves one request or a hundred, so servers batch many requests together.

Engines such as vLLM manage the cache in pages to fit more requests in memory, an idea introduced in the PagedAttention paper, and schedule requests with continuous batching. To judge whether those choices help, measure time to first token, time per output token, and throughput together. LLM inference metrics defines each one.

Where AI inference runs

Common places to run inference and what each trades off.
OptionWhat you controlTypical fit
Hosted model APIPrompts and parameters onlyFast start, variable or low traffic
Cloud GPUs you rentModel, engine, and scalingSteady traffic, custom or open models, data control
Your own data centerEverything, including hardwareVery large, predictable workloads
On-device or edgeA small model on a phone, laptop, or embedded chipOffline use, privacy, very low latency

Most teams that serve open models use cloud GPUs. AI cloud infrastructure for LLM inference explains how that stack fits together, and GPU memory for LLM inference shows how to pick a GPU for a given model.

How inference is optimized

Inference optimization means producing the same quality of answer with less time, memory, or money. The most widely used techniques are:

  • Quantization: storing weights in fewer bits so they take less memory and are faster to read.
  • Speculative decoding: letting a cheap draft propose several tokens that the main model checks in one step.
  • Continuous batching: adding and removing requests from the running batch at every step.
  • Prefix caching: reusing the KV cache of a shared system prompt or document across requests. See vLLM vs SGLang for how two engines approach it.

Each of these trades something: quantization can affect accuracy, and larger batches can slow individual streams. Change one at a time, measure, and check answer quality alongside speed. The vLLM optimization guide lists the settings that control these tradeoffs in one engine.

AI inference FAQ

Is inference cheaper than training? Per request, yes, by a wide margin. In total it can cost more, because a popular model answers requests every second for as long as it is deployed.

Does a model learn during inference? No. The weights are fixed. A chat model “remembers” earlier messages only because they are sent again as part of the input.

Do I need a GPU for inference? Not always. Small models run on CPUs and phones. Large language models are usually served on GPUs or other accelerators because they need high memory bandwidth.

What is an inference engineer? An engineer who makes trained models fast, reliable, and affordable to serve. How to learn inference engineering lays out the skills step by step, and the free Foundations course is a good place to start.

Sources and further reading

Primary documentation and research behind this guide.

  1. Attention Is All You Need (2017)
  2. Hugging Face Transformers: generation with LLMs
  3. Efficient Memory Management with PagedAttention (2023)
  4. vLLM: optimization and tuning

Keep learning

Related guideLLM inference explained: from prompt to generated tokensRelated guideGPU memory for LLM inference: how to size and choose a GPURelated guideAI cloud infrastructure for LLM inference: a practical guideRelated guideHow to learn inference engineering: a step-by-step roadmapAcademy courseInference Engineering FoundationsAcademy coursevLLM
Back to all articles