AI Fundamentals
Embeddings Explained: The Building Block Behind Semantic Search
Stop thinking about search as keyword matching. Understand embeddings as numerical representations of meaning that enable true semantic retrieval.
Imagine you are building a support center for a product. A user types into your search bar:
“How do I get my money back?”
Your traditional keyword search looks for the words “money” and “back”. But your actual help document says:
“Customers can request a refund within 30 days of purchase.”
The words are different. There is no overlap between “money back” and “refund”. In a keyword-based world, this is a miss. In a semantic world, it’s a perfect match.
This gap is where embeddings come in.
Try to Understand It This Way
Think of an embedding as a learned numerical representation of text that places semantically related inputs closer together in a vector space.
Imagine we had a magical map of all human concepts. On this map, “Apple” (the fruit) would be physically close to “Pear” and “Banana,” but very far from “Airplane” or “Justice.”
An embedding is simply a set of coordinates on that map.
If two sentences have similar meanings, their coordinates will be close to each other. If they are unrelated, they will be far apart.
Crucially, this map doesn’t care about the specific words used. It cares about the concept the words represent.
What an Embedding Actually Looks Like
In a codebase, an embedding isn’t a map or a concept. It’s just an array of floating-point numbers.
[0.012, -0.042, 0.731, -0.112, 0.004, ...]
A typical embedding from a model like OpenAI’s text-embedding-3-small contains 1,536 dimensions.
A common mistake is trying to interpret these numbers. You might wonder, “Does the 12th number represent ‘financial intent’ or the 400th represent ‘politeness’?”
The answer is: no.
Embeddings are distributed representations learned by the model. The “meaning” is distributed across all dimensions. Individual numbers are not usually meaningful to humans; what matters is the relationship between vectors when we compare them using a similarity measure.
From Text to Vectors: The Pipeline
The process of turning text into an embedding is straightforward:
"How do I get a refund?"
↓
Embedding Model
↓
[0.12, -0.42, 0.73, ...]
You send a string to an embedding model, and it returns a vector. This vector is a compressed representation of the semantic features of that text.
Measuring Meaning: Cosine Similarity
Once we have these vectors, how do we know if two pieces of text are “similar”?
We use math. The most common method in AI engineering is Cosine Similarity.
Instead of measuring the distance between the tips of two vectors (like a ruler), cosine similarity measures the angle between them.
- A score of
1.0means the vectors point in the same direction. - A score of
0.0means the vectors are orthogonal (at a right angle). - A score of
-1.0means they point in opposite directions.
Mathematically, cosine similarity ranges from -1 to 1. In most semantic search tasks, we primarily care about the range between 0 and 1. The higher the score, the more likely the vectors represent related concepts, though the exact semantic meaning of a score depends on the model and the data.
Under the Hood: The Math of Cosine Similarity
If you have two vectors $\mathbf{A}$ and $\mathbf{B}$, the cosine similarity is the dot product of the vectors divided by the product of their magnitudes:
$$\text{similarity} = \frac{\sum_{i=1}^{n} A_i B_i}{\sqrt{\sum_{i=1}^{n} A_i^2} \sqrt{\sum_{i=1}^{n} B_i^2}}$$
In simpler terms:
- Multiply the corresponding numbers in each vector and sum them up (Dot Product).
- Divide by the “length” of each vector to normalize them.
This ensures that the length of the text doesn’t artificially inflate the similarity score.
The Semantic Search Flow
When you build a semantic search system, you aren’t searching the text. You are searching the vectors.
Here is the complete engineering flow:
-
Index Time (Prepare the data):
- Take your documents.
- Pass each one through the embedding model.
- Store the resulting vectors (and the original text).
-
Query Time (Handle the user):
- Take the user’s query.
- Pass it through the same embedding model.
- Compare the query vector against all stored document vectors using cosine similarity.
- Rank the documents by the highest score.
Critical Rule: Use the same embedding model and compatible configuration for documents and queries so they live in the same vector space. If you embed your documents with OpenAI and your query with Ollama, the coordinates come from different embedding spaces, so their similarity scores are not directly comparable.
In practice: embed your documents and queries with the same embedding model/configuration.
Cross-Language Implementation
One of the most important realizations for a new AI engineer is that embeddings are a data concept, not a language concept.
Whether you are using Python, JavaScript, Rust, or Go, the pipeline is identical:
API Call $\rightarrow$ Vector $\rightarrow$ Math.
Try the Demo
The concepts are easier to understand when you see the complete flow working end to end.
I’ve implemented this small semantic-search example in both Python and JavaScript. The demo keeps things intentionally simple: generate embeddings, calculate cosine similarity, and rank documents based on semantic similarity — without introducing a vector database.
You can explore the implementations here:
- JavaScript: Embedding Semantic Search — JavaScript
- Python: Embedding Semantic Search — Python
The goal is not to build a production-ready search system, but to make the underlying mechanics visible. Once you understand this basic flow, the next step is understanding how vector indexes and vector databases make the same retrieval process work efficiently at much larger scale.
Engineering Decisions: The Trade-offs
When moving this from a demo to production, several decisions matter.
1. Model Choice
A larger or more capable embedding model may improve retrieval quality for some workloads, but it can also increase cost, latency, or vector-storage requirements. Benchmark the model on your own retrieval task rather than assuming bigger is always better.
2. Dimensions
Some modern models support configurable dimensions. This allows you to reduce the size of the vector (e.g., from 1536 to 256 dimensions) to reduce storage costs and speed up search. However, reducing dimensions is a model-specific capability, and retrieval quality can change. The right choice should be evaluated against your actual workload.
3. Normalization
Whether embeddings are normalized depends on the model and provider. Check the model documentation before assuming unit-length vectors.
When vectors are normalized, cosine similarity is equivalent to their dot product, which is computationally cheaper.
4. Embedding Costs
Embedding generation is not just a query-time concern. For a single user query, you generate one vector. But at scale, index-time costs become significant: if you have 1 million documents, you must generate (and potentially re-generate) 1 million embeddings. Batching, caching, and index-update strategies are critical production concerns.
What Breaks?
Semantic search is powerful, but it is not a silver bullet.
The “Related but Different” Problem
Embeddings measure similarity, not correctness.
If a user searches for “How to cancel my subscription,” a document about “How to start a subscription” might have a very high similarity score because both discuss “subscriptions” and “account management,” even though the intent is opposite.
Semantic similarity alone is often not enough when your application requires precise distinctions between related but opposite actions.
Ambiguous Queries
Short queries like “Payment” are ambiguous. Does the user want to know how to pay, why a payment failed, or how to change their payment method? An ambiguous query can produce an embedding that does not strongly represent any one intended meaning, which can lead to less precise retrieval.
Domain-Specific Terminology
General-purpose models are trained on the broad internet. If your company uses highly specific internal jargon (e.g., “Project X-15 Blue-Phase”), a general embedding model may not understand the nuance. In those cases, you may need a domain-specific model or a hybrid search approach.
When I Would Use It
I lean toward embeddings-based search when:
- The users’ vocabulary differs from the documentation’s vocabulary.
- I need to find “similar” items (recommendations).
- I want to support multiple languages without translating everything first (using multilingual embeddings).
- I’m building the first layer of a RAG pipeline.
Things to Think About
To sharpen your intuition, try to imagine how this would behave in these scenarios:
- The Scale Problem: This demo calculates similarity in-memory for five documents. You can compare a query against every vector, but doing so becomes increasingly expensive as the dataset grows. Large-scale systems use indexing and approximate nearest-neighbor (ANN) retrieval to avoid scanning every vector for every query. This is the core problem Article #4 solves. Ask GPT ↗️
- The Model Drift: What happens to your indexed vectors if the embedding provider updates their model version? Do you have to re-embed everything? Ask GPT ↗️
- The Precision Gap: What if two documents are semantically similar, but only one is the legally correct answer for the user? How do you handle that? Ask GPT ↗️
- The Multi-lingual Shift: If you embed a Spanish query and an English document using a multilingual model, will they match? Why? Ask GPT ↗️
Key Takeaways
- Embeddings are learned numerical representations of text.
- Semantic Search compares vectors, not keywords.
- Cosine Similarity is a standard way to measure the angle between vectors.
- Same Model Always: Use the same embedding model and configuration for documents and queries to ensure they live in the same vector space.
Related Sources
- OpenAI Embeddings Documentation — The official guide to OpenAI’s embedding models.
- Ollama Documentation — For running embedding models locally.
- Cosine Similarity (Wikipedia) — The mathematical foundation of vector comparison.
Have a question? Talk to GPT
If you want to explore these concepts further, try these prompts:
- Why is cosine similarity preferred over Euclidean distance for text embeddings? Ask GPT ↗️
- How do “dimensions” in an embedding actually represent features of the text? Ask GPT ↗️
- When would a keyword search actually be better than a semantic search? Ask GPT ↗️
What’s Next?
This works beautifully for a small collection of documents. But as an engineer, you know that “in-memory” doesn’t scale.
If you have millions of vectors, comparing against every single one for every request becomes increasingly expensive. You need specialized indexing and retrieval techniques to find the nearest vectors efficiently.
In the next article, we’ll explore Vector Databases Explained: What Problem Are They Solving?