An embedding model turns text into a vector so that similar meanings land close together. Contextator runs one locally on your CPU by default — no API key, no account, nothing leaving the machine — and can use OpenAI instead if you prefer.
The default
| Model | Xenova/multilingual-e5-small |
| Dimensions | 384 |
| Languages | 100, including Turkish. Ask in the language the answer is written in: matching across languages is a measured limit of the model, not a switch |
| Download | ~470 MB once (fp32), cached in a Docker volume |
| Runs on | CPU. No GPU is used or needed |
Smaller and faster options
EMBEDDING_DTYPE=q8 # same model, ~120 MB, quantized
EMBEDDING_MODEL=Xenova/all-MiniLM-L6-v2 # English-only, ~90 MB, also 384 dimensions
EMBEDDING_DIMENSIONS=384
Any transformers.js feature-extraction model works. EMBEDDING_DIMENSIONS must match what the model
produces — if it does not, the server tells you the correct number.
Using OpenAI instead
EMBEDDING_PROVIDER=openai
OPENAI_API_KEY=sk-…
OPENAI_EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_DIMENSIONS=1536
RESET_VECTORS=1 # only for the first start after the change — see below
Better retrieval on long chunks and a much larger input window, at the cost of an API call per chunk indexed and per query asked. Your documentation text is sent to OpenAI; if that is not acceptable, stay local.
Changing the model
Same dimensions (e.g. between the 384-dimension local models)
- Change
EMBEDDING_MODEL. - Restart.
- Re-index each project.
Contextator records which model indexed each project. When the server’s model differs, the next run is
automatically a full re-index, and search_docs refuses to search that project until then — with a
message saying exactly that, rather than returning confidently wrong results.
Different dimensions (e.g. moving to OpenAI’s 1536)
The vector column’s width is fixed when the database is created, so this is a one-time destructive step:
- Set the provider, key and
EMBEDDING_DIMENSIONS. - Start once with
RESET_VECTORS=1. Every chunk and document is dropped and the column is re-typed. - Remove
RESET_VECTORS(or set it back to0). - Re-index every project.
Without the flag, the server refuses to start and prints this instruction:
The database was created with EMBEDDING_DIMENSIONS=384 but the current config says 1536.
Either set EMBEDDING_DIMENSIONS=384, or start once with RESET_VECTORS=1
(drops every indexed chunk; all projects must be re-indexed), or wipe the pgdata volume.
How search actually works
Search is hybrid: meaning and exact wording are searched at once and fused by rank.
- Your question is embedded with the same model that indexed the documents.
- PostgreSQL runs two rankings scoped to that project — nearest chunks by cosine similarity over an
HNSW index, and a
tsvectorkeyword match over a GIN index — and combines them with reciprocal rank fusion, so an environment variable finds its page as readily as a question does. - The top matches are returned with their file path, heading breadcrumb and score.
Scores are cosine similarities: higher is better, and with the default model they sit high and close
together. On the shipped configuration the mean score is 0.881, and a wrong hit sits only about 0.02
below a correct one. A score is therefore not a truth signal, and nothing in the product
reads one to decide anything.
The relevance floor, SEARCH_SCORE_FLOOR=0.82 by default, was measured against something else
entirely: questions with nothing to do with your documentation, whose top hits reach only 0.829,
where the lowest top hit of a real question is 0.833. That gap is the whole of what the floor claims.
What it does not catch is a question shaped like your product whose answer is simply not written
down — those score exactly where real questions score, and no threshold separates them. Below the floor
search_docs answers no good match rather than handing over a hit an agent would cite.
Scores are comparable within one project and one model, and meaningless across different models — change
EMBEDDING_MODEL and the floor has to be re-measured with npm run eval.
SEARCH_SCORE_FLOOR (default 0.82) is the server’s floor. A project can set its own from the query
panel, or turn the floor off for itself, when its corpus scores differently — prose-heavy corpora usually
want a lower one. The panel shows what a floor would have done to the searches already logged before it
is applied, and the server’s SEARCH_SCORE_FLOOR=0 still turns every project’s floor off. See
Configuration.
Ranks are fused, not scores. Cosine similarity and PostgreSQL’s ts_rank are not comparable
quantities, so the fusion combines the rank each chunk held on each side — Σ 1/(60 + rank) over
whichever lists it appeared on (reciprocal rank fusion) — rather than a weighted sum that would have to
be re-learned every time the embedding model changed. The consequence: a result’s score no longer
explains its position — one scoring 0.86 can sit above one scoring 0.88. GET /api/projects/:id/search
and the dashboard show the rank each result held on each side as D3 L1; a result with an L and no D
is an identifier the vector search alone could not see. On the golden set, rank fusion took the thirty
identifier-shaped questions from recall@5 76.7 % to 83.3 % and did not move the thirty-four
natural-language questions at all; recall@1 fell from 78.1 % to 75.0 %, the trade fusion makes — a
chunk found by one half alone cannot outrank one found respectably by both.
The default model is asymmetric: it was trained with query: in front of a search query and
passage: in front of an indexed passage, so the server adds them automatically — nothing you index or
search ever mentions them, and a model without prefixes encodes a query and a passage identically at no
cost. On the retrieval corpus the prefixes were worth one question at rank 1 and nothing at recall@5;
EMBEDDING_QUERY_PREFIX=none with EMBEDDING_PASSAGE_PREFIX=none turns them off and restores the
previous model id exactly.
The keyword half indexes each source in PostgreSQL’s simple configuration by default — words as
written, no stemming — which is the right choice for reference material: HALYARD_DISPATCH_TIMEOUT and
HLY-4015 survive into the index as the strings they are. A source can instead name its language (the
Language field on the source form), which switches it to a stemmed configuration — a Turkish question
asking about anahtarı then finds a page that says anahtarın, two strings simple treats as unrelated.
The two live in one project without interfering: each source is searched in its own configuration and a
single query reaches all of them, each contributing its own ranked list to the fusion. Changing a
source’s language re-indexes it.
Cross-lingual search is a known limit, and it is not being fixed
Write the question in the same language as the documentation that should answer it. Turkish pages
surface for phrasing shaped like Turkish. English pages surface for phrasing shaped like English.
Crossing that boundary mostly returns nothing, and this is a property of the embedding model rather than
a setting to turn on. Every MCP client is a language model, and list_topics names a source’s language
while instructions directs the calling agent to write its search_docs query to match — a direction to
the caller, not a change to search itself.
The measurement is in eval/BASELINE.md
and it is not close: thirty questions were written from the corpus, each about a page in the language the
question was not phrased in — fifteen per direction. Four are answered in the top five. Twenty-seven of
the thirty instead return a page in the phrasing’s own language at rank 1: the model is not failing to
understand what was asked, it is ranking the language of the phrasing above the answer to it. Both
directions fail equally.
Identifiers are the exception. HALYARD_DISPATCH_TIMEOUT, HLY-4015, X-Halyard-Signature are the
same string in both languages, so the keyword half of search finds them regardless. Of the nine
cross-lingual questions in the set that name an identifier, four are answered; of the twenty-one phrased
as ordinary questions, none is.
What to do instead of waiting for a fix. Keep each language in its own source and let an agent scope
its search with source, path_prefix or version — two sources that each answer well beat one
collection that answers either language badly. Both known remedies have already been measured: hybrid
search (above) recovered cross-lingual identifier retrieval and no natural-language question; a
multilingual cross-encoder rerank behind SEARCH_RERANK took cross-lingual recall@5 from 13.3 % to
33.3 % — still below the 42.9 % the embedding model before this one managed — while costing thirteen
ordinary questions their rank-1 answer and a search’s latency going from 12 ms to 1.2 s. It is off by
default and should stay off. A second, translation-trained encoder as an additional index would actually
fix this — roughly four times the download, a second vector column, a full re-index on every
installation — and that price is not worth paying for a documentation server. If cross-language search
is essential to your corpus, that is a reason to choose a different tool rather than a reason to wait for
this one.
Air-gapped installations
The only outbound request a default installation makes is the one-time model download. To remove even that:
- Pre-populate the models volume (
/app/.cache/models) with the model files — for example by running once on a connected host and copying the volume. - Set
EMBEDDING_OFFLINE=1.
The server will then never attempt a download, and will fail loudly if the model is missing.
Performance notes
- Embedding is CPU-bound and is the slow part of indexing; that is why one project is indexed at a time.
EMBEDDING_BATCH_SIZE(default 16) controls how many chunks are embedded per call.- A query embeds exactly one short text, so search latency is dominated by the database, not the model.
- The model is loaded in the background at startup — the dashboard works while it loads, but indexing and search wait for it.