Everything is an environment variable. With Compose, put them in .env next to docker-compose.yml;
the shipped .env.example is the
annotated full list.
Configuration is validated at startup. If something is wrong, the server prints every problem and exits rather than starting in a half-working state.
Server
| Variable | Default | What it does |
|---|---|---|
PORT |
3444 |
The port the application listens on |
HOST |
0.0.0.0 |
Bind address inside the container. Leave this alone — it is not the setting that controls what is reachable from outside; that is CONTEXTATOR_BIND below. Docker cannot forward a published port to a process bound only to the container’s own loopback |
CONTEXTATOR_BIND |
127.0.0.1 |
Which host interface docker compose publishes the port on (Compose only — the application itself never reads it). The default answers on this machine and nowhere else. 0.0.0.0 publishes on every interface — do that only with a reverse proxy or a VPN in front, and set TRUST_PROXY and PUBLIC_BASE_URL to match; see Running behind a reverse proxy below |
LOG_LEVEL |
info |
fatal…trace. Per-request logging turns on at debug |
PUBLIC_BASE_URL |
– | e.g. https://docs.example.com. Used for the URLs shown in the dashboard, and — behind a proxy — for the OAuth metadata MCP connectors read. See Running behind a reverse proxy |
TRUST_PROXY |
0 |
Which peers may tell this server where a request came from. Decides req.ip (the per-IP sign-in limit, /oauth/register’s per-host budget, the address beside an audit event) and req.protocol. 0 ignores X-Forwarded-* entirely, which is correct for the shipped Compose shape with nothing in front of it. Behind a real proxy this must be set — see Running behind a reverse proxy below |
ALLOWED_ORIGINS |
– | Comma-separated browser origins allowed to call /mcp/*. Non-browser clients are always allowed |
SESSION_IDLE_TTL_MS |
1800000 |
Idle Streamable HTTP transport sessions are closed after 30 minutes. Dashboard sign-ins are separate — that is AUTH_SESSION_IDLE_MS |
MCP_STRUCTURED_OUTPUT |
0 |
1 makes the MCP tools publish an outputSchema and return structuredContent beside the unchanged text. Off by default because Claude Code reads the structured content instead of the text once a result carries it — see MCP Tools |
Accounts and sessions
People sign in to the dashboard with their own account; these settings shape how that works. See Accounts and Permissions.
| Variable | Default | What it does |
|---|---|---|
SETUP_CODE |
– | The one-time code /setup asks for before the first account exists. Set it and you never have to read it out of the log; leave it empty and the server generates one and prints it at every start until that account is created. Ignored from then on. Case, dashes and punctuation are ignored when it is checked, so give it enough letters and digits rather than relying on punctuation |
ADMIN_TOKEN |
– | Machine access to /api/* via Authorization: Bearer …, acting with root permissions — for scripts and CI. Browsers sign in with an account instead; treat this like a root password and never paste it into one. Does not protect /mcp/* |
AUTH_SESSION_IDLE_MS |
43200000 (12 h) |
A dashboard session unused for this long has to sign in again. Refreshed while the dashboard is in use |
AUTH_SESSION_TTL_DAYS |
30 |
Hard ceiling on a session’s life, however actively it is used. Must not be shorter than AUTH_SESSION_IDLE_MS |
AUTH_COOKIE_SECURE |
auto |
auto sets the Secure flag when the request arrives over HTTPS. Force it with 1 behind a terminating proxy; use 0 for a plain-HTTP LAN install, or the browser drops the cookie and sign-in loops back to /login |
AUTH_LOGIN_MAX_ATTEMPTS |
10 |
Failed sign-ins per account and per IP address before a lockout or a 429 |
AUTH_LOGIN_WINDOW_MIN |
15 |
The IP window in minutes, and the first step of the account lockout — which doubles from there, capped at an hour |
PASSWORD_MIN_LENGTH |
12 |
Applies to every password, temporary ones included. No composition rules are imposed |
MCP_OAUTH |
1 |
The OAuth 2.1 flow that lets a client act as an account on /mcp/*. 0 removes it, and then only static tokens open a closed project — browser-based connectors cannot connect at all |
MCP_OAUTH_ACCESS_TTL_MIN |
60 |
Life of one access token handed to a connector |
MCP_OAUTH_REFRESH_TTL_DAYS |
30 |
A connector left unused this long has to be authorized again |
Federated sign-in (OIDC_ISSUER_URL and the rest of the OIDC_* variables) is its own page — see
Single Sign-On.
Database
| Variable | Default | What it does |
|---|---|---|
POSTGRES_PASSWORD |
contextator |
Password of the embedded PostgreSQL. Applied when the cluster is first created; changing it later needs ALTER USER — see Backup and Data |
DATABASE_URL |
– | Local development only. The container ignores it and talks to its embedded PostgreSQL |
RESET_VECTORS |
0 |
One-time destructive reset when changing embedding dimensions — see Embedding Models |
Storage (Compose only)
| Variable | Default | What it does |
|---|---|---|
CONTEXTATOR_PGDATA_VOLUME |
contextator-pgdata |
Volume name for the database |
CONTEXTATOR_MODELS_VOLUME |
contextator-models |
Volume name for downloaded models |
CONTEXTATOR_DATA_VOLUME |
contextator-data |
Volume name for materialised sources |
CONTEXTATOR_PGDATA_PATH |
– | Absolute host directory used instead of the volume above |
CONTEXTATOR_MODELS_PATH |
– | Same, for models |
CONTEXTATOR_DATA_PATH |
– | Same, for materialised sources |
DOCS_HOST_PATH |
./docs |
Host folder mounted read-only at /docs |
Examples:
CONTEXTATOR_PGDATA_PATH=/srv/contextator/pgdata # Linux: keep the database on a chosen disk
CONTEXTATOR_MODELS_PATH=D:/contextator/models # Windows host directory (forward slashes)
CONTEXTATOR_PGDATA_VOLUME=contextator-pgdata-v2 # or simply another named volume
Documents
| Variable | Default | What it does |
|---|---|---|
ALLOWED_DOC_ROOTS |
/docs |
Comma-separated. A local directory source must live inside one of these. This is a real security boundary — see Security |
IGNORE_GLOBS |
– | Comma-separated globs skipped while indexing, e.g. **/CHANGELOG.md,drafts/**. Applies to every source |
DATA_DIR |
.data (/data in the container) |
Writable directory holding git checkouts, uploads and Notion pulls |
SECRET_KEY |
– | At least 32 characters (openssl rand -hex 32). Encrypts git, Notion and Confluence tokens and webhook secrets at rest (AES-256-GCM). Needed only once such a source exists. Changing it on its own leaves every stored token unreadable — rotate it instead: set SECRET_KEY_PREVIOUS to the retiring key, set the new SECRET_KEY, restart, run npm run rotate-secret, then remove SECRET_KEY_PREVIOUS and restart again. See the rotation runbook on Security |
SECRET_KEY_PREVIOUS |
– | The key being retired, set only for the length of a rotation (above). It never encrypts — reads fall back to it, every write uses SECRET_KEY |
CONFLUENCE_ALLOWED_HOSTS |
– | Comma-separated host names (or IP literals) a Confluence source may reach on a private address (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, fc00::/7), e.g. wiki.corp.example. Needed for a Data Center instance on the internal network; loopback, link-local (169.254.169.254 included), unspecified and multicast addresses stay refused whatever is listed. Checked on the connected address and again on every redirect (ADR-0088). See Confluence |
MAX_STORED_DOCUMENT_BYTES |
1048576 (1 MB) |
How much of each document’s text is kept in the database for read_document. Past it, the prefix is stored and the tool says so. Compressed out of line by PostgreSQL, so the cost is a small fraction of the same document’s vectors |
MAX_CONVERTED_FILE_BYTES |
33554432 (32 MiB) |
Checked against the size the scan recorded, before the file is read, so a limit is never applied to bytes already in memory. .md, .mdx and .txt are decoded rather than parsed and are not capped |
MAX_PDF_PAGES |
2000 |
Refuses a PDF that declares more pages than this before every page is read into memory at once |
MAX_DOCX_UNPACKED_BYTES |
268435456 (256 MiB) |
A .docx is a zip whose directory reports a size its author writes; each part is inflated through a counter and discarded once this ceiling is crossed, which is what actually bounds a zip bomb |
CONVERSION_TIMEOUT_MS |
120000 (2 min) |
A file still converting past this is refused by name, the same as an unreadable PDF; the worker thread is replaced and the run carries on |
CONVERSION_IDLE_MS |
60000 (1 min) |
The conversion worker thread is kept warm between files and dropped after this much idle time, so a quiet server is not holding the heap a large document grew |
Uploads
| Variable | Default | What it does |
|---|---|---|
UPLOAD_MAX_FILE_BYTES |
52428800 (50 MB) |
Per uploaded file |
UPLOAD_MAX_FILES_PER_REQUEST |
500 |
The dashboard splits large folders across requests by itself |
UPLOAD_MAX_ARCHIVE_BYTES |
268435456 (256 MB) |
Per uploaded archive |
ARCHIVE_MAX_ENTRIES |
20000 |
Guard applied while extracting |
ARCHIVE_MAX_TOTAL_BYTES |
1073741824 (1 GB) |
Guard applied while extracting |
Embeddings
| Variable | Default | What it does |
|---|---|---|
EMBEDDING_PROVIDER |
local |
local (CPU, no API key) or openai |
EMBEDDING_MODEL |
Xenova/multilingual-e5-small |
Any transformers.js feature-extraction model. Previous default, reads only 128 tokens: Xenova/paraphrase-multilingual-MiniLM-L12-v2. English-only and faster: Xenova/all-MiniLM-L6-v2. All three are 384-d |
EMBEDDING_DIMENSIONS |
384 |
Must match the model. 1536 for text-embedding-3-small |
EMBEDDING_DTYPE |
fp32 |
q8 downloads a ~4× smaller quantized model |
EMBEDDING_BATCH_SIZE |
16 |
Chunks embedded per batch |
MODEL_CACHE_DIR |
.cache/models |
/app/.cache/models inside the container |
EMBEDDING_OFFLINE |
0 |
1 forbids model downloads (air-gapped hosts with a pre-populated cache) |
OPENAI_API_KEY |
– | Required when the provider is openai |
OPENAI_EMBEDDING_MODEL |
text-embedding-3-small |
|
EMBEDDING_MAX_INPUT_TOKENS |
– | What the model reads usefully — the window it was trained at, not where its tokenizer truncates. Left empty, the server discovers it from the loaded model and warns after startup if CHUNK_MAX_TOKENS does not fit; set, it overrules that discovery, and a contradicting CHUNK_MAX_TOKENS refuses to start |
EMBEDDING_QUERY_PREFIX / EMBEDDING_PASSAGE_PREFIX |
– | The instruction prefixes the model was trained with, put on by the server and never by you. Empty means the model decides: query: / passage: for multilingual-e5-*, nothing for anything else. The trailing space matters and .env strips an unquoted one, so write EMBEDDING_QUERY_PREFIX="query: ". An empty value reads as unset rather than no prefix — none is how you say no prefix on a model that has them. Either value is part of the model id: changing one re-indexes every project, exactly as changing EMBEDDING_MODEL does. See Embedding Models |
SEARCH_SCORE_FLOOR |
0.82 |
Cosine similarity below which search_docs answers no good match rather than its best hit. 0 turns it off. Measured against the default embedding model and meaningless on another one |
Search and ranking
Tuning for the vector index and how one answer is assembled from its hits. The defaults are what the golden set measured best; see How Indexing Works for the benchmark these came from.
| Variable | Default | What it does |
|---|---|---|
HNSW_EF_SEARCH |
100 |
How many candidates the vector index produces before the generation filter is applied — pgvector’s own default is 40. Every project has its own partial index, so the candidates are that project’s rows; while a project re-indexes, its next generation sits in the same index and is filtered out afterwards. Costs latency on every search |
HNSW_ITERATIVE_SCAN |
relaxed_order |
relaxed_order, strict_order or off. Keeps scanning when the filter leaves fewer hits than asked for, instead of answering short (pgvector 0.8+; on an older release all three settings are ignored and search behaves as it did before). relaxed_order returns the rows unordered and the server sorts them itself |
HNSW_MAX_SCAN_TUPLES |
20000 |
The ceiling that actually ends an iterative scan, counted in tuples of the project’s own index. Raise it as a project grows, not as the instance does |
SEARCH_MAX_PER_DOCUMENT |
2 |
Excerpts one document may contribute to one answer, applied after ranking and refilled from the excerpts below it, so an agent that asked for five still gets five. Measured on the golden set it gains a question — what it drops is a near-duplicate of something already on the page. 20 turns it off |
SEARCH_NEIGHBOR_CONTEXT |
1 |
Chunks either side of each hit, shown as context around it rather than as further results. 0 turns it off. A chunk is CHUNK_MAX_TOKENS, so one either side is about three times the context a hit used to be |
SEARCH_MAX_RESULT_CHARS |
12000 |
Ceiling on one rendered search_docs answer; past it whole excerpts are dropped and the result says how many. A default answer is around 3 300 characters |
Sync
| Variable | Default | What it does |
|---|---|---|
SYNC_DEFAULT_INTERVAL_MINUTES |
60 |
The sync interval a newly created source is given, in minutes; 0 creates it unscheduled. It never reaches a source that already exists — not on upgrade, and not when this value changes — so an upgrade starts no outbound traffic nobody asked for. Per source, the dashboard and the API accept 5 to 43200 (30 days), or never |
SYNC_PROBES_PER_TICK |
10 |
How many due sources one tick — one minute — may check. The rest keep their turn, oldest first, and the next tick takes them |
Observability
| Variable | Default | What it does |
|---|---|---|
AUDIT_LOG_RETENTION_DAYS |
365 |
How long an audit event is kept. There is no switch for the log itself — every state-changing admin request that succeeds is recorded. See Admin API |
METRICS_TOKEN |
– | A bearer credential that reaches GET /metrics and nothing else, so scraping does not mean handing Prometheus an ADMIN_TOKEN. At least 16 characters — the length is a floor, not entropy, and this endpoint is not rate-limited. Unset, /metrics still answers a signed-in account or ADMIN_TOKEN, but only while the database is up; an instance that wants to stay readable during an outage sets this. See Admin API |
METRICS_PUBLIC |
0 |
1 answers /metrics with no credential at all. For a private network, or a proxy that already guards the path — anywhere the port is reachable, leave it off; the exposition describes the instance |
Chunking
| Variable | Default | What it does |
|---|---|---|
CHUNK_MAX_TOKENS |
96 |
Tokens per chunk, counted with the embedding model’s own tokenizer |
CHUNK_OVERLAP_TOKENS |
24 |
Overlap between consecutive chunks of one section. Must be smaller than CHUNK_MAX_TOKENS |
Worth knowing:
96is not a fraction of the model’s window — the default model reads 512 tokens — it is what measured best on the golden set, and filling the window measures worse. Raise both on OpenAI, whose window is 8191. Changing either only affects files indexed afterwards — force a re-index to apply it everywhere. See Indexing.
Running behind a reverse proxy
Nothing is in front of this by default — Compose publishes 3444 on the loopback interface
(CONTEXTATOR_BIND=127.0.0.1), with nothing between it and the port, and the defaults above are written
for that. Publishing it anywhere else (CONTEXTATOR_BIND=0.0.0.0) is the moment to put something in
front. Put nginx, Caddy, Traefik or a cloud load balancer in front and two settings have to move
together:
TRUST_PROXY=172.18.0.0/16 # the proxy's own address or CIDR, or `loopback` if it runs on the host
PUBLIC_BASE_URL=https://docs.example.com
AUTH_COOKIE_SECURE=1
TRUST_PROXY is what lets this server read X-Forwarded-For and X-Forwarded-Proto from that proxy.
Left at 0 behind one, three things break, and none of them says so on its own:
- The per-IP sign-in limit becomes instance-wide. Every request carries the proxy’s address, so
AUTH_LOGIN_MAX_ATTEMPTSfailures from one person answer429to everybody. - MCP connectors stop being able to authorize. With
PUBLIC_BASE_URLunset, the OAuth protected-resource metadata and theWWW-Authenticatepointer are built fromreq.protocol— which readshttp— so the document advertiseshttp://…while the client sendshttps://…, and/oauth/authorizeanswersinvalid_targetevery time. SettingPUBLIC_BASE_URLfixes this half on its own, which is why it is in the block above. - The session cookie loses its
Secureflag, becauseAUTH_COOKIE_SECURE=autofollows the same forwarded scheme.AUTH_COOKIE_SECURE=1fixes this half on its own.
The server watches for the mistake rather than guessing at it: the first request that arrives carrying
an X-Forwarded-* header this instance is not trusting logs one warning naming all three.
Name the proxy, not the network your clients are on. The list is matched against every hop, not
only against the peer — the server walks the chain from the socket outwards, and req.ip is the first
address the list does not cover. So TRUST_PROXY=uniquelocal on a LAN where the clients are also on
10.0.0.0/8 or 192.168.0.0/16 protects nothing: a client is walked past exactly as a proxy is, and
X-Forwarded-For: 203.0.113.99 puts that value into req.ip again. TRUST_PROXY=1 is the same hazard
stated plainly, and is only safe when nothing but the proxy can reach the port at all. A hop count is
not accepted — it is a claim the server cannot check, and it goes silently wrong the day a CDN
appears in front of the proxy.
Make sure the proxy replaces X-Forwarded-For rather than appending to whatever the client sent
(nginx: proxy_set_header X-Forwarded-For $remote_addr;, not $proxy_add_x_forwarded_for, unless a
further trusted proxy sits in front of it). A proxy that appends hands the caller the left-most value,
and no setting here can tell the difference. See Security
for an nginx example that puts Basic authentication in front of the dashboard.
Applying changes
docker compose up -d # recreates the container with the new .env
Most settings take effect immediately. A few need a little more:
EMBEDDING_MODEL— the next index run becomes a full re-index automatically. Nothing starts that run by itself: the project page shows the mismatch with a Re-index now button, and search stays refused until it is pressed. The run is safe to start at any time of day — it is written beside the old index and published in one step at the end, so the project never passes through “no indexed content”. While it runs, the project holds two copies of its chunks and its share of the vector index, so plan disk for the peak. See Embedding Models.EMBEDDING_QUERY_PREFIX/EMBEDDING_PASSAGE_PREFIX— count as changing the model: setting either, including setting both tonone, makes every project re-index on its next run.EMBEDDING_DIMENSIONS— requiresRESET_VECTORS=1once; see Embedding Models.CHUNK_MAX_TOKENS/CHUNK_OVERLAP_TOKENS— apply to newly indexed files; use Force re-index to rebuild everything.
Upgrading across a default change is its own case: Xenova/multilingual-e5-small became the default
after Xenova/paraphrase-multilingual-MiniLM-L12-v2. An installation that never set EMBEDDING_MODEL
picks the new one up on upgrade, and every existing project reads as a mismatch until it is re-indexed.
Pin the old value in .env to postpone that — and note that going back is not a matter of reverting the
setting alone: a project already re-indexed under the new model needs another re-index to return.
Checking what is live
GET /api/health reports the running configuration: database status, embedding provider, model,
dimensions, dtype and readiness and its input window, whether CHUNK_MAX_TOKENS fits that window, open
MCP sessions, allowed document roots, the data directory, whether SECRET_KEY is configured, and the
upload limits. The dashboard’s top bar shows the important ones.
It answers with less detail the less it knows about the caller: an anonymous request — a monitor, or
the container’s own health check — gets ok, the version, the database status and whether setup is
still pending, and nothing that describes the machine. Sign in, or use ADMIN_TOKEN, for the rest.
GET /metrics is a separate, Prometheus-shaped endpoint with its own credential rules — see
Admin API.