AI Fundamentals

Tokens, Context Windows, and Why They Matter

Understand what an LLM actually sees, how text becomes tokens, and why context windows become an engineering constraint in production AI systems.

October 7, 2026

In the previous article, we built a mental model for an LLM as a probabilistic prediction engine.

That immediately raises a practical question:

What exactly does the prediction engine see when we give it a prompt?

If you send an LLM this:

“What is our refund policy?”

it doesn’t receive the sentence as a human would read it.

It receives a sequence of tokens.

And once you start building real AI applications, tokens stop being an implementation detail.

They affect:

  • how much information you can send to a model
  • how much a request costs
  • how much context the model can work with
  • how long a request can take
  • what information gets dropped when a conversation becomes too large
  • how you design RAG, agents, chat history, and long-running workflows

This is why understanding tokens and context windows is not just an LLM fundamentals exercise.

It is an engineering problem.

Start With a Simple Question: What Is a Token?

Let’s start with something familiar.

Take this sentence:

“The refund policy is available in our customer portal.”

A human sees words.

An LLM tokenizer breaks the text into tokens.

A token is a piece of text that the model processes as an input unit. Depending on the tokenizer and the text, a token might represent a whole word, part of a word, punctuation, whitespace, or another common piece of text.

For example, a word such as:

refund

might be represented as one token in one tokenizer.

A less common or longer word might be split into several pieces.

The exact tokenization depends on the model and its tokenizer.

So the useful mental model is not:

one word = one token

It is:

text → tokens → model

That distinction matters because the model’s limits and usage are generally measured in tokens, not words.

Why Don’t Models Just Use Words?

Because language has too many possible words and word forms to make a simple word-based representation practical.

Consider these:

connect
connected
connecting
connection
reconnection

A tokenizer can often represent these using reusable pieces rather than requiring every possible word form to be treated as completely unrelated.

The same idea becomes even more useful for names, technical terms, URLs, code, numbers, and unfamiliar words.

A tokenizer is essentially creating a vocabulary of reusable pieces that the model can process.

You don’t need to memorize how tokenization works to build AI applications.

But you do need to understand one consequence:

The amount of text you send is not the same thing as the number of tokens you consume.

Try to Understand It This Way

Imagine you are packing a suitcase.

You don’t measure the suitcase by the number of shirts.

You measure the space each item actually occupies.

Two shirts might take roughly the same amount of space.

But a pair of shoes takes more.

A jacket takes more.

A strange-shaped object might take even more.

Tokens work in a similar way.

Two sentences with the same number of words can consume different numbers of tokens.

That is why engineers working with LLMs eventually stop thinking only in terms of:

“How many words are in this prompt?”

and start thinking:

“How many tokens are going into this request?”

This becomes especially important when your application starts sending large amounts of context.

The Context Window

Now we can introduce the second important concept:

the context window.

Think of the context window as the amount of tokenized information the model can consider within a single request.

It includes more than just the user’s latest message.

Depending on the application, the context might contain:

  • system instructions
  • conversation history
  • the user’s current message
  • retrieved documents
  • tool results
  • previous workflow state
  • examples included in the prompt
  • other application-generated context

Imagine a customer-support application.

The user asks:

“Can I get a refund?”

Your application might construct something closer to:

System instructions
+
Conversation history
+
Customer information
+
Refund policy retrieved from knowledge base
+
Current user question

The model doesn’t magically know which of these pieces your application considers important.

Your application decides what goes into the context.

That is one of the most important ideas to understand:

The context window is not just a limit. It is part of the interface between your application and the model.

A Conversation Is Not the Same Thing as Memory

This distinction becomes important very quickly.

Imagine you are building a chatbot.

The conversation starts:

User: I have a problem with my order.

Then:

User: It was order #48291.

Then:

User: It arrived yesterday.

Then:

User: The product is damaged.

When the user eventually asks:

“Can I get a replacement?”

the model needs enough relevant information from the previous conversation to answer correctly.

But the model does not automatically get an infinite transcript.

Your application has to decide what conversation state should be included in the request.

That could mean sending:

  • the full conversation
  • a summarized conversation
  • selected relevant messages
  • structured application state
  • some combination of these

This is why conversation history and model memory are not the same thing.

Your application owns the state.

The model receives whatever context your application chooses to provide.

What Happens When the Context Gets Too Large?

This is where the concept becomes an engineering problem.

Imagine a conversation that has grown to hundreds of messages.

You also retrieve several documents for every question.

You add system instructions.

You add tool results.

Eventually, the total context becomes too large for the model’s available context window.

You now have a design problem.

Something has to change.

Your application might:

  1. Remove older messages.
  2. Summarize earlier conversation.
  3. Retrieve only the most relevant information.
  4. Reduce the amount of retrieved content.
  5. Store important information as structured state instead of raw conversation.
  6. Split a long workflow into multiple model calls.

This is why simply choosing a model with a larger context window doesn’t eliminate the problem.

A larger window gives you more room.

It doesn’t tell you what deserves that room.

The Hidden Cost of “Just Send Everything”

One of the easiest mistakes to make in an early LLM application is this:

“The model supports a large context window, so let’s just send everything.”

Imagine a support conversation containing 100 messages.

Now imagine your application retrieves 20 documents for every new question.

Then your prompt contains:

System instructions
+
100 messages
+
20 documents
+
tool results
+
current question

The model might technically be able to process it.

But that doesn’t automatically make it a good system.

More context can mean:

  • more tokens
  • higher cost
  • more latency
  • more information competing for the model’s attention
  • more irrelevant content
  • more complicated debugging
  • potentially worse answers when important information is buried in noise

The engineering question becomes:

What is the smallest useful context that gives the model enough information to do the job correctly?

That question will come up again and again throughout this series.

Tokens Affect Cost

Most LLM APIs charge based, directly or indirectly, on token usage.

A request can contain input tokens and produce output tokens.

So consider two versions of the same application.

Version A

System instructions:     500 tokens
Conversation:           2,000 tokens
Retrieved context:      1,000 tokens
User question:            50 tokens
---------------------------------
Input:                  3,550 tokens

Version B

System instructions:     500 tokens
Conversation:           8,000 tokens
Retrieved context:      6,000 tokens
User question:            50 tokens
---------------------------------
Input:                 14,550 tokens

Both applications might answer the same question.

But Version B is sending roughly four times as much input context.

At scale, that difference matters.

If the application handles thousands or millions of requests, inefficient context construction can become a significant cost driver.

This is one reason techniques such as:

  • retrieval
  • chunking
  • summarization
  • caching
  • context filtering
  • prompt optimization

matter in production AI systems.

We will explore several of these later in the series.

Tokens Affect Latency Too

Tokens aren’t only about cost.

They can also affect latency.

Before a model can generate an answer, it has to process the input context.

A request with a tiny prompt and a request with a huge prompt are not equivalent workloads.

Then the model has to generate the output itself.

So, at a high level, an LLM request has two different token-related concerns:

Input processing

The model processes the context you send.

Output generation

The model generates the response token by token.

This distinction becomes especially important for applications that stream responses to users.

A longer context can increase the work required before generation begins, while a longer output increases the time spent generating the response.

The exact performance characteristics depend on the model and serving system, but the engineering principle is simple:

Context size is part of your latency budget.

A Production Example: Customer Support

Let’s return to the support bot from Article #1.

A customer asks:

“Can I get a refund for this purchase?”

Your first implementation might retrieve ten documents:

Refund Policy
Shipping Policy
Cancellation Policy
Warranty Policy
Privacy Policy
Account Policy
...

Then you place all of them into the prompt.

It works.

So you might think:

“Great. The model has all the information.”

But the better question is:

“Did we give the model the information it actually needs?”

If the answer only depends on the refund policy, sending nine unrelated documents is unnecessary context.

A better architecture might:

  1. Identify the user’s intent.
  2. Retrieve relevant refund-policy content.
  3. Include only the useful sections.
  4. Pass that focused context to the model.
  5. Validate the generated answer where required.

The model hasn’t become smarter.

The application has become better at constructing context.

That is a very important production-AI skill.

What About Long Documents?

Now imagine your customer-support team has a 200-page policy document.

You can’t necessarily throw the entire document into every request just because the model has a large context window.

Instead, you might:

  1. Split the document into smaller chunks.
  2. Store those chunks in a searchable system.
  3. Retrieve the relevant chunks for a user’s question.
  4. Put only those chunks into the model’s context.

This is one of the foundations of Retrieval-Augmented Generation (RAG).

We’ll spend much more time on RAG later in this series.

For now, the important idea is:

Don’t confuse a large context window with a good retrieval strategy.

A large context window gives you capacity.

Retrieval gives you selectivity.

Context Is an Engineering Resource

Once you start thinking in tokens, context begins to look like another engineering resource.

We already think about:

  • CPU
  • memory
  • network bandwidth
  • database connections
  • queue capacity

Now add:

  • context budget

You have a finite amount of useful context you can send to a model.

And you need to decide how to spend it.

For example:

Context budget
│
├── System instructions
├── Conversation state
├── Retrieved knowledge
├── Tool results
└── Current request

Every additional piece of context has a cost.

The question isn’t:

“Can I fit this?”

The better question is:

“Is this information valuable enough to deserve space in the context?”

That is the mental model I want you to carry forward.

What Changed for Me as an Engineer

One of the things that changed for me when I started building AI systems was realizing that context construction is part of application logic.

In traditional backend systems, we are used to explicitly deciding which database records, API responses, or configuration values a piece of code needs.

With LLM applications, we make a similar decision — but the dependency is often expressed as context.

What do we retrieve?

What do we include?

What do we summarize?

What do we remove?

What should become structured state instead of raw conversation?

Those decisions can have as much impact on the behavior of the application as the prompt itself.

That’s why I don’t think of context management as just “prompt engineering.”

It is part of system design.

Practical Rules for the AI Engineer

If you’re starting to build LLM applications, keep these rules in mind:

  1. Think in tokens, not words.

    Your model limits, costs, and request sizes are based on tokens.

  2. Treat context as a budget.

    Don’t send everything simply because you can.

  3. Keep relevant context and remove noise.

    More context is not automatically better context.

  4. Separate conversation history from application state.

    Important information may deserve structured storage rather than being buried inside a growing transcript.

  5. Measure context size in production.

    Track input tokens, output tokens, latency, and cost. You can’t optimize what you don’t observe.

Try This Mental Exercise

Imagine you’re building an AI assistant for an insurance company.

A user asks:

“Am I covered if my car is damaged by flooding?”

Your application has access to:

  • the user’s policy
  • a 200-page product document
  • the claims process
  • previous conversation history
  • a knowledge base
  • internal tools

Would you send everything to the model?

Probably not.

You would want to identify the relevant information, retrieve it, structure the context, and then ask the model to generate the answer.

Now change the question:

“What is the status of my claim?”

The relevant context is different.

You probably need application state or a tool call rather than a large collection of policy documents.

Same model.

Different context.

Different application behavior.

That’s the key lesson.

Think About These Use Cases

Try these questions to sharpen your intuition:

  • Long conversations: At what point would you stop sending the entire chat history and start summarizing or selecting relevant messages?
    Ask GPT ↗️

  • RAG: If your retrieval system returns 20 chunks but only three actually answer the question, what should happen to the other 17?
    Ask GPT ↗️

  • Agents: If an agent runs ten tools and each tool returns a large response, how would you prevent the context from growing uncontrollably?
    Ask GPT ↗️

  • Cost: If two applications produce the same answer but one sends four times as many input tokens, what would you investigate first?
    Ask GPT ↗️

  • Debugging: If an LLM suddenly starts giving worse answers, how would you determine whether the problem is the model, the prompt, or the context being supplied?
    Ask GPT ↗️

Further Reading & Sources

  • OpenAI Tokenizer — A useful way to visualize how text is broken into tokens.
  • OpenAI Documentation — Documentation covering tokens, context windows, and model usage.
  • Attention Is All You Need — The research paper that introduced the Transformer architecture and the attention mechanism behind modern LLMs.

Have a question? Talk to GPT

If you want to explore these ideas interactively, try one of these questions:

  • Why are tokens a better engineering unit than words when working with LLMs? Ask GPT ↗️

  • Why doesn’t a larger context window automatically mean a better AI application? Ask GPT ↗️

  • What’s the difference between conversation history, application state, and model context? Ask GPT ↗️

  • How would you design a production system that keeps LLM context small without losing important information? Ask GPT ↗️

What’s Next?

We now know that an LLM works with tokens and that the context we provide is a finite engineering resource.

But there is another question.

If we want an AI system to find information based on meaning rather than exact keywords, how do we represent that meaning in a form a machine can search?

That’s where embeddings come in.

In the next article, we’ll look at Embeddings Explained: The Building Block Behind Semantic Search.