Which Embedding Model Should Power Your Site Search?
Which embedding model should power your site search?
For most B2B sites, the smallest current model from whichever provider you already use. The published quality gap between small and large embedding models is a couple of points on a benchmark. The cost gap is several times. Your search quality will be decided by chunking and content, not by the model.
That is not the answer teams expect. The choice gets framed as picking a winner, so it attracts more debate than it deserves while the parts that actually control relevance get none.
Here is what the vendor documentation actually publishes, what the numbers mean, and where we would spend the attention instead.
What does an embedding model actually do here?
It turns a piece of text into a list of numbers so that similar meanings land near each other. Your search then stops matching words and starts matching meaning, which is why someone can type "make my site load faster" and reach a page titled "Core Web Vitals".
The list of numbers is the embedding and its length is the dimension count. You embed every chunk of your content once, store the vectors, then embed the visitor's query at search time and find the nearest stored vectors.
Everything else in the stack is plumbing around that one idea. We walked through the whole build in our piece on adding AI site search to a website.
How many dimensions do you actually need?
Fewer than the defaults, usually. Both major providers now let you ask for shorter vectors from the same model, and both say the shorter version keeps most of its usefulness.
OpenAI documents this as a dimensions parameter, explaining that "developers can shorten embeddings (i.e. remove some numbers from the end of the sequence) without the embedding losing its concept-representing properties". Its own example is striking: a text-embedding-3-large embedding "can be shortened to a size of 256 while still outperforming an unshortened text-embedding-ada-002 embedding with a size of 1536".
Google describes the same technique by name. Its embeddings documentation says its models use Matryoshka Representation Learning, which "teaches a model to learn high-dimensional embeddings that have initial segments (or prefixes) which are also useful, simpler versions" of the data, and recommends 768, 1536 or 3072 output dimensions.
What do the vendors actually publish?
Enough to compare on paper. These figures are each provider's own published numbers, read on their documentation on 29 September 2026, not independent testing.
| Model | Default dimensions | Max input tokens | Provider's published MTEB score | Provider's stated cost |
|---|---|---|---|---|
| text-embedding-3-small | 1536 | 8192 | 62.3 percent | About 62,500 pages per dollar |
| text-embedding-3-large | 3072 | 8192 | 64.6 percent | About 9,615 pages per dollar |
| gemini-embedding-2 | 3072, configurable from 128 | 8192 | Not on the page we read | Not on the page we read |
| gemini-embedding-001 | 3072, configurable from 128 | 2048 | Not on the page we read | Not on the page we read |
Two things jump out. The quality difference OpenAI publishes between its small and large model is 2.3 points. The cost difference it publishes is roughly six and a half times more pages per dollar for the small one.
The other is the token limit. Google's documentation gives gemini-embedding-2 an 8,192 token input limit and gemini-embedding-001 a 2,048 token limit. If you are embedding long pages without splitting them, that gap matters more than any benchmark score.
Does a higher benchmark score mean better site search?
Not reliably. MTEB is a broad benchmark covering many tasks across many datasets. Your site search is one task on one dataset: your content, your visitors' phrasing, your product vocabulary.
A model that scores two points higher on an average of dozens of tasks may or may not be better at matching "SOC 2" to your security page. The only test that settles it is running both against a list of real queries from your own site search logs.
We think that test is worth an hour and almost nobody runs it. Twenty real queries with a human judging the top three results will tell you more than any leaderboard, and it will usually tell you the models are close and your content is the problem.
What breaks first, the model or the chunking?
The chunking, almost every time. If you embed a 3,000 word page as one vector, you get one blurry average of everything on it. The page matches nothing well because it is about eight things at once.
Split the same page by section and each vector becomes a sharp answer to one question. Search quality jumps, and it jumps whichever model you picked. This is the single highest leverage decision in the whole build.
The practical rule we use is to chunk on headings, keep the heading text inside the chunk, and never let a chunk cross a topic boundary. We went deeper on the retrieval side of this in how page chunking affects AI retrieval.
Where does your database put a ceiling on the choice?
At the index, not the column, and this catches people out. pgvector, the extension most Postgres-based stacks use, states that "vectors can have up to 16,000 dimensions". That sounds like no limit at all.
The indexes are stricter. pgvector documents both HNSW and IVFFlat as supporting "vector - up to 2,000 dimensions". A 3072 dimension embedding fits in the column and cannot be indexed as a plain vector, which means an exact scan of every row at query time.
There are ways round it. pgvector's halfvec type raises the index limit to 4,000 dimensions at half the storage, and its published storage figures are 4 bytes per dimension plus 8 for vector against 2 bytes per dimension plus 8 for halfvec. Or you ask the model for 1536 dimensions and stay inside the plain limit. Either way, decide this before you embed a few hundred thousand chunks.
Which distance measure should you use?
Cosine distance for text search, in almost every case. pgvector supports six operators: L2 distance, negative inner product, cosine distance, taxicab distance, and Hamming and Jaccard distance for binary vectors.
Cosine compares direction rather than magnitude, which is what you want when one chunk is a short heading and another is four paragraphs. Inner product is faster when your vectors are already normalised, and Google notes that its default 3072 dimension embeddings "are always normalized", so that route is open if you measure a benefit.
The index tuning knobs matter more than the operator choice. pgvector's HNSW defaults are m at 16 connections per layer and ef_construction at 64. Raising them costs build time and memory and buys recall. Start with the defaults and only move them with a query set to measure against.
Should you match the task type to the query?
If your provider supports it, yes, because it is free accuracy. Google's gemini-embedding-001 accepts eight task types, including RETRIEVAL_DOCUMENT, RETRIEVAL_QUERY, SEMANTIC_SIMILARITY, CLASSIFICATION, CLUSTERING, CODE_RETRIEVAL_QUERY, QUESTION_ANSWERING and FACT_VERIFICATION.
The pairing that matters for site search is embedding your content as RETRIEVAL_DOCUMENT and your visitor's query as RETRIEVAL_QUERY. The same text embedded under different task types produces different vectors, tuned for different jobs.
Google's documentation notes the newer model handles this differently: for gemini-embedding-2, task types are not set by parameter, and you "include the task as an instruction in your prompt" for text-only tasks. Worth checking which behaviour your chosen model expects before you build around it.
What happens when you need to change models later?
You re-embed everything. Vectors from different models are not comparable, so there is no gradual migration. Whatever you pick, plan for a full rebuild of the index as a routine operation rather than a crisis.
That means keeping the source text and the chunk boundaries as your system of record, with the vectors treated as derived data you can regenerate. If your only copy of a chunk is inside the vector store, a model change becomes an archaeology project.
Providers do retire models, and an automation that silently depends on one is a scheduled outage. We wrote about handling that in what to do when a model you depend on is deprecated.
What would we actually pick, and how will that age?
For a B2B marketing site on Postgres: the smaller current model from the provider already in the stack, asked for 1536 dimensions so it indexes cleanly, cosine distance, heading-based chunking, and a twenty query test set kept in the repository. That configuration is cheap, indexable and good enough that content quality becomes the limiting factor, which is where it should be.
We expect the model choice to keep getting less important. Dimension flexibility, longer input limits and near-identical benchmark scores all push in the same direction: the differentiator moves from which model you chose to how well your content was written and split.
If you are standing up search or retrieval over your own content and want a second opinion on the shape before you commit, we are happy to look at it. Reach us at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.