AI Fundamentals

What Is an LLM and How Does It Actually Work?

Move past the 'magic box' misconception and understand LLMs as probabilistic prediction engines through the lens of an engineer.

October 7, 2026

When I first started working with Large Language Models (LLMs), I fell into a common trap. I treated them like a very sophisticated version of a database.

I thought: “I’ll give it the information, and it will retrieve the correct answer for me.”

But if you’ve spent any time building AI features, you’ve likely realized that LLMs don’t actually retrieve information in the way a SQL query or a search engine does.

In fact, if you treat an LLM like a database, you’ll spend most of your time fighting hallucinations and wondering why the system sometimes behaves unpredictably.

To build reliable AI systems, we have to move away from the “knowledge base” mental model and start thinking of LLMs as probabilistic prediction engines.

The “Autocomplete” Intuition

The most intuitive way to understand an LLM is to look at the autocomplete feature on your phone.

When you type “How are…”, your phone suggests “you”. It isn’t searching a database of every conversation ever had to find the correct completion. It has learned patterns in language and uses the context available to predict what is likely to come next.

An LLM is essentially that same basic mechanism, but scaled up enormously.

Imagine you’re finishing a sentence in a conversation. Someone says:

“I usually drink coffee in the…”

Depending on the context, you might expect “morning”, “afternoon”, or even “office”.

The prediction changes with the context.

This is the core mental model for how an LLM generates text. It doesn’t retrieve a record from a database and return it. It repeatedly predicts what token should come next based on the context it has available.

From Intuition to Engineering: Next-Token Prediction

In technical terms, the primary mechanism here is next-token prediction.

An LLM doesn’t see text exactly the way we do. It processes text as tokens (which we’ll dive into in the next article). Given the tokens in a prompt, the model produces a probability distribution over possible next tokens.

If the prompt is:

“The capital of France is…”

the model’s internal state will assign different probabilities to possible next tokens. “Paris” will generally be much more likely than an unrelated token.

The model then selects a token according to its generation strategy and appends it to the sequence. It then processes the expanded context and predicts the next token again.

This loop continues until the model produces a stop condition or reaches a configured length limit.

At the generation level, the model is repeatedly estimating:

“Given the context so far, what token should come next?”

That sounds simple. But when you combine this mechanism with billions of learned parameters and a large amount of contextual information, surprisingly complex behavior can emerge — including behavior that looks a lot like reasoning.

Want to go deeper? Under the hood of the "Engine"

If you’re wondering how the model makes these predictions, two concepts are particularly useful for building an intuition: weights and attention.

Weights (The Parameters): Think of these as the learned parameters of the neural network. During training, the model processes enormous amounts of text and adjusts these parameters so that its predictions become better over time. The resulting parameters encode statistical patterns and relationships learned from the training data.

Attention (The Contextual Focus): Attention is a core part of the Transformer architecture. It allows the model to determine which parts of the current context are relevant to each token it is processing.

For example, in:

“The bank was closed because the river overflowed.”

the representation of “bank” can attend strongly to words such as “river” and “overflowed”, helping the model interpret “bank” in the geographic sense rather than the financial one.

The Engineering Consequence: Why the “Magic” Breaks

When you treat a prediction engine as a database, your system will eventually run into problems.

Here are two of the most important ones.

1. Hallucinations are plausible continuations

A hallucination isn’t simply a “lie.” The model is not inherently performing a fact-checking operation before producing every statement.

It is generating a plausible continuation from the context and patterns available to it.

That means a model can produce an answer that sounds extremely confident and reasonable while still being factually wrong.

This is why simply telling a model “don’t hallucinate” is not a sufficient engineering solution.

Instead, reliable systems give the model better grounding: relevant context, retrieval, tools, validation, structured outputs, evaluation, or human review depending on the use case.

2. The challenge of probabilistic generation

LLM generation often uses probabilistic sampling. That means the same prompt can produce different outputs across runs when the generation configuration allows multiple plausible choices.

Lowering temperature can make generation more consistent, but production behavior can also depend on the model version, sampling configuration, and serving environment.

For a software engineer, this creates an important shift.

Traditional software often gives us deterministic behavior that we can test with assertions such as:

A -> B

LLM applications require another layer of testing: evaluations.

Instead of asking only:

“Did the function return exactly B?”

we may also need to ask:

“Did the model produce an acceptable answer across a representative set of inputs?”

That is a different engineering problem.

Practical Rules for the AI Engineer

If you’re building production AI systems, I suggest adopting these three rules of thumb.

  1. Don’t rely on model weights as your application’s source of truth.

    Never assume the model “knows” your company’s latest pricing, policies, or a user’s specific account details. Supply critical information through context, retrieval, application state, tools, or another explicit source of truth.

  2. Steer the probability, don’t just give instructions.

    Instead of saying “be concise,” provide examples of the exact format you want when appropriate. With few-shot prompting, you are showing the model the pattern you want it to follow.

  3. Assume the prediction will eventually fail.

    Since generation is probabilistic, your system needs to handle incorrect or unexpected outputs. Your job is to build guardrails — validation, structured outputs, retries, evaluation, and human review where necessary — that prevent model failures from becoming user-facing failures.

Case Study: The Customer Support Bot

Imagine we are building a support bot for a fintech app.

A customer asks:

“What is your refund policy?”

If we rely on the model’s internal knowledge to answer that question, it might generate a generic, plausible-sounding refund policy based on patterns it learned during training.

That is a high-risk failure.

The engineering approach is different:

  1. Retrieve the actual refund policy from the application’s approved knowledge source.
  2. Inject the relevant information into the model’s context.
  3. Instruct the model to answer using the provided policy.
  4. Validate the response where the workflow requires it.
  5. Keep a human in the loop when the risk or business process requires human approval.

Now the model is no longer being asked to invent a policy from what it learned during training. It is generating an answer grounded in information supplied by the application.

The model is still a prediction engine.

We’ve simply designed the surrounding system so that it has better information to predict from.

That distinction becomes increasingly important as AI systems move from demos into production.

What changed for me as an engineer

Moving from traditional software to AI engineering required me to accept a different kind of uncertainty.

In traditional software, when a deterministic function produces the wrong result, I can usually trace the failure to code, data, configuration, or an external dependency.

With LLM applications, another source of failure exists: the model can produce an output that is syntactically valid, plausible, and still wrong.

I’ve learned that the model is just one component in a larger system.

The real engineering happens in the orchestration around it: how we retrieve information, how we provide context, how we validate outputs, how we evaluate behavior, and how we observe the system in production.

That is one of the biggest mental shifts when moving from building software that follows explicit rules to building software that works with probabilistic models.

Think about these use cases

To sharpen your intuition, try to imagine how a “prediction engine” would handle these scenarios compared to a “database”.

I’ve included some guided questions you can explore with GPT:

  • A Customer Support Bot: If a user asks about a refund policy, does the model “know” the policy, or is it generating an answer based on the policy text you provided in its context? Ask GPT ↗️

  • Code Generation: When an LLM writes a Python function, is it “reasoning” through the logic, or is it generating a likely sequence of tokens based on patterns it learned during training? Where can this distinction cause a bug? Ask GPT ↗️

  • Creative Writing: Why is a probabilistic prediction engine useful for generating a poem when a deterministic database is not? Ask GPT ↗️

Further Reading & Sources

  • Attention Is All You Need — The original research paper that introduced the Transformer architecture.
  • OpenAI Documentation — Documentation and guides covering model behavior and building with OpenAI models.
  • Andrej Karpathy’s “Let’s build GPT” — A hands-on technical walkthrough for engineers who want to understand how a GPT-style model can be built from the ground up.

Have a question? Talk to GPT

If you want to dive deeper into any of these concepts, you can ask GPT directly. I’ve pre-written a few prompts to get you started.

  • Why does the “prediction engine” model explain hallucinations better than the “database” model? Ask GPT ↗️

  • How does “Temperature” actually affect the probability distribution of the next token? Ask GPT ↗️

  • Can you give me a real-world example of where treating an LLM as a database would lead to a production failure? Ask GPT ↗️

  • If LLMs are just predicting the next token, how are they able to perform complex reasoning tasks? Ask GPT ↗️

What’s next?

Now that we understand the “prediction engine” mental model, we need to look at the fuel that powers it.

In the next article, we’ll explore Tokens and Context Windows — and why the way an LLM processes text is one of the first major constraints you’ll encounter when building AI systems in production.