Contextator
ENTR

Operating

Embedding Models

The default local model, switching to OpenAI, changing models safely, how search actually scores results, and running fully air-gapped.

Updated:

An embedding model turns text into a vector so that similar meanings land close together. Contextator runs one locally on your CPU by default — no API key, no account, nothing leaving the machine — and can use OpenAI instead if you prefer.

The default

Model Xenova/multilingual-e5-small
Dimensions 384
Languages 100, including Turkish. Ask in the language the answer is written in: matching across languages is a measured limit of the model, not a switch
Download ~470 MB once (fp32), cached in a Docker volume
Runs on CPU. No GPU is used or needed

Smaller and faster options

EMBEDDING_DTYPE=q8                       # same model, ~120 MB, quantized
EMBEDDING_MODEL=Xenova/all-MiniLM-L6-v2  # English-only, ~90 MB, also 384 dimensions
EMBEDDING_DIMENSIONS=384

Any transformers.js feature-extraction model works. EMBEDDING_DIMENSIONS must match what the model produces — if it does not, the server tells you the correct number.

Using OpenAI instead

EMBEDDING_PROVIDER=openai
OPENAI_API_KEY=sk-…
OPENAI_EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_DIMENSIONS=1536
RESET_VECTORS=1        # only for the first start after the change — see below

Better retrieval on long chunks and a much larger input window, at the cost of an API call per chunk indexed and per query asked. Your documentation text is sent to OpenAI; if that is not acceptable, stay local.


Changing the model

Same dimensions (e.g. between the 384-dimension local models)

  1. Change EMBEDDING_MODEL.
  2. Restart.
  3. Re-index each project.

Contextator records which model indexed each project. When the server’s model differs, the next run is automatically a full re-index, and search_docs refuses to search that project until then — with a message saying exactly that, rather than returning confidently wrong results.

Different dimensions (e.g. moving to OpenAI’s 1536)

The vector column’s width is fixed when the database is created, so this is a one-time destructive step:

  1. Set the provider, key and EMBEDDING_DIMENSIONS.
  2. Start once with RESET_VECTORS=1. Every chunk and document is dropped and the column is re-typed.
  3. Remove RESET_VECTORS (or set it back to 0).
  4. Re-index every project.

Without the flag, the server refuses to start and prints this instruction:

The database was created with EMBEDDING_DIMENSIONS=384 but the current config says 1536.
Either set EMBEDDING_DIMENSIONS=384, or start once with RESET_VECTORS=1
(drops every indexed chunk; all projects must be re-indexed), or wipe the pgdata volume.

How search actually works

Search is hybrid: meaning and exact wording are searched at once and fused by rank.

  1. Your question is embedded with the same model that indexed the documents.
  2. PostgreSQL runs two rankings scoped to that project — nearest chunks by cosine similarity over an HNSW index, and a tsvector keyword match over a GIN index — and combines them with reciprocal rank fusion, so an environment variable finds its page as readily as a question does.
  3. The top matches are returned with their file path, heading breadcrumb and score.

Scores are cosine similarities: higher is better, and with the default model they sit high and close together. On the shipped configuration the mean score is 0.881, and a wrong hit sits only about 0.02 below a correct one. A score is therefore not a truth signal, and nothing in the product reads one to decide anything.

The relevance floor, SEARCH_SCORE_FLOOR=0.82 by default, was measured against something else entirely: questions with nothing to do with your documentation, whose top hits reach only 0.829, where the lowest top hit of a real question is 0.833. That gap is the whole of what the floor claims. What it does not catch is a question shaped like your product whose answer is simply not written down — those score exactly where real questions score, and no threshold separates them. Below the floor search_docs answers no good match rather than handing over a hit an agent would cite.

Scores are comparable within one project and one model, and meaningless across different models — change EMBEDDING_MODEL and the floor has to be re-measured with npm run eval.

SEARCH_SCORE_FLOOR (default 0.82) is the server’s floor. A project can set its own from the query panel, or turn the floor off for itself, when its corpus scores differently — prose-heavy corpora usually want a lower one. The panel shows what a floor would have done to the searches already logged before it is applied, and the server’s SEARCH_SCORE_FLOOR=0 still turns every project’s floor off. See Configuration.

Ranks are fused, not scores. Cosine similarity and PostgreSQL’s ts_rank are not comparable quantities, so the fusion combines the rank each chunk held on each side — Σ 1/(60 + rank) over whichever lists it appeared on (reciprocal rank fusion) — rather than a weighted sum that would have to be re-learned every time the embedding model changed. The consequence: a result’s score no longer explains its position — one scoring 0.86 can sit above one scoring 0.88. GET /api/projects/:id/search and the dashboard show the rank each result held on each side as D3 L1; a result with an L and no D is an identifier the vector search alone could not see. On the golden set, rank fusion took the thirty identifier-shaped questions from recall@5 76.7 % to 83.3 % and did not move the thirty-four natural-language questions at all; recall@1 fell from 78.1 % to 75.0 %, the trade fusion makes — a chunk found by one half alone cannot outrank one found respectably by both.

The default model is asymmetric: it was trained with query: in front of a search query and passage: in front of an indexed passage, so the server adds them automatically — nothing you index or search ever mentions them, and a model without prefixes encodes a query and a passage identically at no cost. On the retrieval corpus the prefixes were worth one question at rank 1 and nothing at recall@5; EMBEDDING_QUERY_PREFIX=none with EMBEDDING_PASSAGE_PREFIX=none turns them off and restores the previous model id exactly.

The keyword half indexes each source in PostgreSQL’s simple configuration by default — words as written, no stemming — which is the right choice for reference material: HALYARD_DISPATCH_TIMEOUT and HLY-4015 survive into the index as the strings they are. A source can instead name its language (the Language field on the source form), which switches it to a stemmed configuration — a Turkish question asking about anahtarı then finds a page that says anahtarın, two strings simple treats as unrelated. The two live in one project without interfering: each source is searched in its own configuration and a single query reaches all of them, each contributing its own ranked list to the fusion. Changing a source’s language re-indexes it.

Cross-lingual search is a known limit, and it is not being fixed

Write the question in the same language as the documentation that should answer it. Turkish pages surface for phrasing shaped like Turkish. English pages surface for phrasing shaped like English. Crossing that boundary mostly returns nothing, and this is a property of the embedding model rather than a setting to turn on. Every MCP client is a language model, and list_topics names a source’s language while instructions directs the calling agent to write its search_docs query to match — a direction to the caller, not a change to search itself.

The measurement is in eval/BASELINE.md and it is not close: thirty questions were written from the corpus, each about a page in the language the question was not phrased in — fifteen per direction. Four are answered in the top five. Twenty-seven of the thirty instead return a page in the phrasing’s own language at rank 1: the model is not failing to understand what was asked, it is ranking the language of the phrasing above the answer to it. Both directions fail equally.

Identifiers are the exception. HALYARD_DISPATCH_TIMEOUT, HLY-4015, X-Halyard-Signature are the same string in both languages, so the keyword half of search finds them regardless. Of the nine cross-lingual questions in the set that name an identifier, four are answered; of the twenty-one phrased as ordinary questions, none is.

What to do instead of waiting for a fix. Keep each language in its own source and let an agent scope its search with source, path_prefix or version — two sources that each answer well beat one collection that answers either language badly. Both known remedies have already been measured: hybrid search (above) recovered cross-lingual identifier retrieval and no natural-language question; a multilingual cross-encoder rerank behind SEARCH_RERANK took cross-lingual recall@5 from 13.3 % to 33.3 % — still below the 42.9 % the embedding model before this one managed — while costing thirteen ordinary questions their rank-1 answer and a search’s latency going from 12 ms to 1.2 s. It is off by default and should stay off. A second, translation-trained encoder as an additional index would actually fix this — roughly four times the download, a second vector column, a full re-index on every installation — and that price is not worth paying for a documentation server. If cross-language search is essential to your corpus, that is a reason to choose a different tool rather than a reason to wait for this one.

Air-gapped installations

The only outbound request a default installation makes is the one-time model download. To remove even that:

  1. Pre-populate the models volume (/app/.cache/models) with the model files — for example by running once on a connected host and copying the volume.
  2. Set EMBEDDING_OFFLINE=1.

The server will then never attempt a download, and will fail loudly if the model is missing.

Performance notes

  • Embedding is CPU-bound and is the slow part of indexing; that is why one project is indexed at a time.
  • EMBEDDING_BATCH_SIZE (default 16) controls how many chunks are embedded per call.
  • A query embeds exactly one short text, so search latency is dominated by the database, not the model.
  • The model is loaded in the background at startup — the dashboard works while it loads, but indexing and search wait for it.

Arrow keys to move, Enter to open.