# Embedding Models

> The default local model, switching to OpenAI, changing models safely, how search actually scores results, and running fully air-gapped.

- Updated: 2026-09-21
- Source: https://contextator.com/en/docs/embedding-models/
- Language: en-US
- Author: Muhammet Şafak

---
An embedding model turns text into a vector so that similar meanings land close together. Contextator
runs one **locally on your CPU** by default — no API key, no account, nothing leaving the machine — and
can use OpenAI instead if you prefer.

## The default

| | |
|---|---|
| Model | `Xenova/multilingual-e5-small` |
| Dimensions | 384 |
| Languages | 100, including Turkish. Ask in the language the answer is written in: matching across languages is a measured limit of the model, not a switch |
| Download | ~470 MB once (`fp32`), cached in a Docker volume |
| Runs on | CPU. No GPU is used or needed |

### Smaller and faster options

```bash
EMBEDDING_DTYPE=q8                       # same model, ~120 MB, quantized
```

```bash
EMBEDDING_MODEL=Xenova/all-MiniLM-L6-v2  # English-only, ~90 MB, also 384 dimensions
EMBEDDING_DIMENSIONS=384
```

Any transformers.js feature-extraction model works. `EMBEDDING_DIMENSIONS` **must** match what the model
produces — if it does not, the server tells you the correct number.

## Using OpenAI instead

```bash
EMBEDDING_PROVIDER=openai
OPENAI_API_KEY=sk-…
OPENAI_EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_DIMENSIONS=1536
RESET_VECTORS=1        # only for the first start after the change — see below
```

Better retrieval on long chunks and a much larger input window, at the cost of an API call per chunk
indexed and per query asked. Your documentation text is sent to OpenAI; if that is not acceptable, stay
local.

---

## Changing the model

### Same dimensions (e.g. between the 384-dimension local models)

1. Change `EMBEDDING_MODEL`.
2. Restart.
3. Re-index each project.

Contextator records which model indexed each project. When the server's model differs, the next run is
automatically a **full** re-index, and `search_docs` refuses to search that project until then — with a
message saying exactly that, rather than returning confidently wrong results.

### Different dimensions (e.g. moving to OpenAI's 1536)

The vector column's width is fixed when the database is created, so this is a one-time destructive step:

1. Set the provider, key and `EMBEDDING_DIMENSIONS`.
2. Start **once** with `RESET_VECTORS=1`. Every chunk and document is dropped and the column is re-typed.
3. Remove `RESET_VECTORS` (or set it back to `0`).
4. Re-index every project.

Without the flag, the server refuses to start and prints this instruction:

```
The database was created with EMBEDDING_DIMENSIONS=384 but the current config says 1536.
Either set EMBEDDING_DIMENSIONS=384, or start once with RESET_VECTORS=1
(drops every indexed chunk; all projects must be re-indexed), or wipe the pgdata volume.
```

---

## How search actually works

Search is **hybrid**: meaning and exact wording are searched at once and fused by rank.

1. Your question is embedded with the same model that indexed the documents.
2. PostgreSQL runs two rankings scoped to that project — nearest chunks by cosine similarity over an
   HNSW index, and a `tsvector` keyword match over a GIN index — and combines them with reciprocal rank
   fusion, so an environment variable finds its page as readily as a question does.
3. The top matches are returned with their file path, heading breadcrumb and score.

Scores are cosine similarities: higher is better, and with the default model they sit high and close
together. On the shipped configuration the mean score is `0.881`, and a **wrong** hit sits only about `0.02`
below a correct one. A score is therefore not a truth signal, and nothing in the product
reads one to decide anything.

The relevance floor, `SEARCH_SCORE_FLOOR=0.82` by default, was measured against something else
entirely: questions with **nothing to do with your documentation**, whose top hits reach only `0.829`,
where the lowest top hit of a real question is `0.833`. That gap is the whole of what the floor claims.
What it does **not** catch is a question shaped like your product whose answer is simply not written
down — those score exactly where real questions score, and no threshold separates them. Below the floor
`search_docs` answers *no good match* rather than handing over a hit an agent would cite.

Scores are comparable within one project and one model, and meaningless across different models — change
`EMBEDDING_MODEL` and the floor has to be re-measured with `npm run eval`.

`SEARCH_SCORE_FLOOR` (default `0.82`) is the server's floor. A project can set its own from the query
panel, or turn the floor off for itself, when its corpus scores differently — prose-heavy corpora usually
want a lower one. The panel shows what a floor would have done to the searches already logged before it
is applied, and the server's `SEARCH_SCORE_FLOOR=0` still turns every project's floor off. See
[Configuration](/en/docs/configuration/).

**Ranks are fused, not scores.** Cosine similarity and PostgreSQL's `ts_rank` are not comparable
quantities, so the fusion combines the **rank** each chunk held on each side — `Σ 1/(60 + rank)` over
whichever lists it appeared on (reciprocal rank fusion) — rather than a weighted sum that would have to
be re-learned every time the embedding model changed. The consequence: **a result's score no longer
explains its position** — one scoring `0.86` can sit above one scoring `0.88`. `GET /api/projects/:id/search`
and the dashboard show the rank each result held on each side as `D3 L1`; a result with an `L` and no `D`
is an identifier the vector search alone could not see. On the golden set, rank fusion took the thirty
identifier-shaped questions from `recall@5` 76.7 % to 83.3 % and did not move the thirty-four
natural-language questions at all; `recall@1` fell from 78.1 % to 75.0 %, the trade fusion makes — a
chunk found by one half alone cannot outrank one found respectably by both.

**The default model is asymmetric**: it was trained with `query: ` in front of a search query and
`passage: ` in front of an indexed passage, so the server adds them automatically — nothing you index or
search ever mentions them, and a model without prefixes encodes a query and a passage identically at no
cost. On the retrieval corpus the prefixes were worth one question at rank 1 and nothing at `recall@5`;
`EMBEDDING_QUERY_PREFIX=none` with `EMBEDDING_PASSAGE_PREFIX=none` turns them off and restores the
previous model id exactly.

**The keyword half indexes each source in PostgreSQL's `simple` configuration by default** — words as
written, no stemming — which is the right choice for reference material: `HALYARD_DISPATCH_TIMEOUT` and
`HLY-4015` survive into the index as the strings they are. A source can instead **name its language** (the
*Language* field on the source form), which switches it to a stemmed configuration — a Turkish question
asking about `anahtarı` then finds a page that says `anahtarın`, two strings `simple` treats as unrelated.
The two live in one project without interfering: each source is searched in its own configuration and a
single query reaches all of them, each contributing its own ranked list to the fusion. Changing a
source's language re-indexes it.

## Cross-lingual search is a known limit, and it is not being fixed

**Write the question in the same language as the documentation that should answer it.** Turkish pages
surface for phrasing shaped like Turkish. English pages surface for phrasing shaped like English.
Crossing that boundary mostly returns nothing, and this is a property of the embedding model rather than
a setting to turn on. Every MCP client is a language model, and `list_topics` names a source's language
while `instructions` directs the calling agent to write its `search_docs` query to match — a direction to
the caller, not a change to search itself.

The measurement is in [`eval/BASELINE.md`](https://github.com/Contextator/Contextator/blob/main/eval/BASELINE.md)
and it is not close: thirty questions were written from the corpus, each about a page in the language the
question was not phrased in — fifteen per direction. Four are answered in the top five. Twenty-seven of
the thirty instead return a page in the phrasing's own language at rank 1: the model is not failing to
understand what was asked, it is ranking the language of the phrasing above the answer to it. Both
directions fail equally.

**Identifiers are the exception.** `HALYARD_DISPATCH_TIMEOUT`, `HLY-4015`, `X-Halyard-Signature` are the
same string in both languages, so the keyword half of search finds them regardless. Of the nine
cross-lingual questions in the set that name an identifier, four are answered; of the twenty-one phrased
as ordinary questions, none is.

**What to do instead of waiting for a fix.** Keep each language in its own source and let an agent scope
its search with `source`, `path_prefix` or `version` — two sources that each answer well beat one
collection that answers either language badly. Both known remedies have already been measured: hybrid
search (above) recovered cross-lingual *identifier* retrieval and no natural-language question; a
multilingual cross-encoder rerank behind `SEARCH_RERANK` took cross-lingual `recall@5` from 13.3 % to
33.3 % — still below the 42.9 % the embedding model before this one managed — while costing thirteen
ordinary questions their rank-1 answer and a search's latency going from 12 ms to 1.2 s. It is off by
default and should stay off. A second, translation-trained encoder as an additional index would actually
fix this — roughly four times the download, a second vector column, a full re-index on every
installation — and that price is not worth paying for a documentation server. If cross-language search
is essential to your corpus, that is a reason to choose a different tool rather than a reason to wait for
this one.

## Air-gapped installations

The only outbound request a default installation makes is the one-time model download. To remove even
that:

1. Pre-populate the models volume (`/app/.cache/models`) with the model files — for example by running
   once on a connected host and copying the volume.
2. Set `EMBEDDING_OFFLINE=1`.

The server will then never attempt a download, and will fail loudly if the model is missing.

## Performance notes

- Embedding is CPU-bound and is the slow part of indexing; that is why one project is indexed at a time.
- `EMBEDDING_BATCH_SIZE` (default 16) controls how many chunks are embedded per call.
- A query embeds exactly one short text, so search latency is dominated by the database, not the model.
- The model is loaded in the background at startup — the dashboard works while it loads, but indexing
  and search wait for it.
