Where everything is stored, what survives what, and how to back it up.
What is stored where
| What | Path in the container | Default volume | Override |
|---|---|---|---|
| Database: projects, documents, embeddings, run history | /var/lib/postgresql/data |
contextator-pgdata |
CONTEXTATOR_PGDATA_VOLUME or CONTEXTATOR_PGDATA_PATH |
| Downloaded embedding models | /app/.cache/models |
contextator-models |
CONTEXTATOR_MODELS_VOLUME / ..._PATH |
| Materialised sources: uploads, git checkouts, Notion pulls | /data |
contextator-data |
CONTEXTATOR_DATA_VOLUME / ..._PATH |
| Your documentation | /docs (read-only) |
– | DOCS_HOST_PATH |
What survives what
| Action | Data |
|---|---|
docker compose down |
Kept |
docker compose pull && docker compose up -d |
Kept |
docker rm contextator, image upgrades |
Kept |
docker compose down -v |
Deleted |
docker volume rm contextator-pgdata |
Deleted |
Your own documentation is never modified: /docs is mounted read-only and local sources are scanned in
place.
On an installation that set DATABASE_URL to an external PostgreSQL, the first row is not yours to back
up from here — the command still runs, dumps the server you named, and says so in its output; that
server’s own backup regime is what actually covers it. On the -slim image, npm run backup refuses
outright, because that image carries no PostgreSQL and therefore no pg_dump: run it from the default
image pointed at the same DATABASE_URL instead.
Storing data in host directories
Instead of named volumes:
CONTEXTATOR_PGDATA_PATH=/srv/contextator/pgdata # Linux: put the database on a chosen disk
CONTEXTATOR_MODELS_PATH=D:/contextator/models # Windows (forward slashes)
CONTEXTATOR_DATA_PATH=/srv/contextator/data
Directories are created on first start and their ownership is fixed automatically. Named volumes remain
the faster choice for the database on Docker Desktop; if initdb reports permission errors on a host
directory, switch that one mount back to a volume.
Backing up
One command, and it takes more than the database:
docker exec contextator npm run backup -- /data/backups/contextator-$(date +%F).tar.gz
docker cp contextator:/data/backups/contextator-$(date +%F).tar.gz .
The command creates the directory it is given. Copy the resulting file off this machine — a backup sitting on the same disk as the thing it backs up is a rollback, not a backup.
What is in the archive
database.dump |
Every project, source, document, chunk, embedding, account, session, MCP token, audit event and logged search — including the encrypted git, Notion and Confluence tokens |
data/… |
The files of every upload source. For uploads this is the only copy in existence; git checkouts and Notion pulls are not in here because they can be re-cloned and re-pulled |
manifest.json |
What the archive is, what is in it, and which SECRET_KEY it needs. It is the first entry, so it can be read without unpacking the rest |
README.txt |
The same thing in prose, for whoever opens the archive in a year |
SECRET_KEY is not in the archive, on purpose
The archive records a fingerprint of SECRET_KEY — a keyed HMAC, not the key itself — never the key.
A backup carrying the key would be the whole instance in one file, which is exactly what encrypting the
source tokens was meant to prevent. Keep .env somewhere the archive is not, and keep a retired key
around if you still hold archives taken before your last SECRET_KEY rotation and those archives carry
a source’s sync credential — a private git, Notion or Confluence token — encrypted under it — see
wiki/Security#rotating-secret_key
for the rotation procedure and why restoring such an archive under the new key refuses.
On a schedule
There is no scheduler inside Contextator — the command is a command, and your host’s cron (or a Compose sidecar) runs it:
#!/usr/bin/env bash
set -euo pipefail
stamp=$(date +%F)
docker exec contextator npm run backup -- "/data/backups/contextator-$stamp.tar.gz"
docker cp "contextator:/data/backups/contextator-$stamp.tar.gz" /srv/backups/
docker exec contextator rm -f "/data/backups/contextator-$stamp.tar.gz"
find /srv/backups -name 'contextator-*.tar.gz' -mtime +14 -delete
set -e on the first line matters: without it, a failed backup is followed by a successful find that
deletes the fortnight it was supposed to replace.
Restoring
docker exec contextator sh -c 'mkdir -p /data/backups'
docker cp contextator-2026-09-22.tar.gz contextator:/data/backups/
# Reads the manifest and every refusal, writes nothing:
docker exec contextator npm run restore -- /data/backups/contextator-2026-09-22.tar.gz --check
docker exec contextator npm run restore -- /data/backups/contextator-2026-09-22.tar.gz
docker compose restart contextator
A restore replaces what is there. It is the right thing on a fresh install and a destructive thing
on a live one, and the command does not ask — --check is how you look first. Before it writes a byte
it confirms the archive’s kind and version, that pg_restore is present, and that the server’s
PostgreSQL major version is not older than the one the dump was taken from — any mismatch there refuses
outright, with an explanation, instead of leaving a half-loaded instance. The SECRET_KEY check is
narrower: it only refuses when the mismatch would lose a sync credential — a git, Notion or
Confluence token nothing on this side can reissue. A webhook secret under the wrong key is not one
of those; it is regenerable, so the restore counts it, reports it, and goes ahead.
A downgrade — restoring a dump from a newer PostgreSQL major into an older server — is refused outright.
Going the other way, 16 to 17, is the documented upgrade path: run npm run backup, start the new major
against an empty volume, then npm run restore the archive into it.
If you only want the database
The two commands underneath are unchanged and still work:
docker exec contextator pg_dump -U contextator -Fc contextator > contextator.dump
docker exec -i contextator pg_restore -U contextator -d contextator --clean --if-exists < contextator.dump
They do not carry the upload files and they do not check SECRET_KEY — which is why npm run backup
and npm run restore exist.
Inspecting the database
The embedded PostgreSQL is not published, so go through the container:
docker exec -it contextator psql -U contextator
\dt -- tables
SELECT name, status, document_count, chunk_count FROM projects;
SELECT name, type, status, document_count FROM document_sources;
SELECT relative_path, title, chunk_count FROM documents ORDER BY relative_path LIMIT 20;
Changing the database password
POSTGRES_PASSWORD is applied only when the cluster is first created. Later:
docker exec -it contextator psql -U contextator -c "ALTER USER contextator PASSWORD 'new-password'"
Then update .env before the next start.
Moving to another machine
npm run backupand copy the archive across.- Copy
.env(includingSECRET_KEY) and your documentation. It must be the key the archive was taken under — if you rotatedSECRET_KEYafter taking it, restore with the retired key instead, or take a fresh backup first. - Start Contextator on the new host, then restore:
mkdir -p /data/backupsin the container,docker cpthe archive in,npm run restore -- <file> --check, thennpm run restore -- <file>, then restart. - Update the URLs your agents point at — or set
PUBLIC_BASE_URLand keep them stable behind a proxy.
Disk usage
- The embedding model: ~470 MB (
fp32) or ~120 MB (q8), once. - The database: dominated by the vectors — 384 floats per chunk plus the chunk text. A few thousand documents is typically in the hundreds of megabytes.
- Materialised sources: as large as the uploads, checkouts and Notion pages themselves.