Contextator
ENTR

Administration & security

Backup and Data

Where Contextator stores its database, models and materialised sources, and how the built-in backup and restore command takes and checks an archive.

Updated:

Where everything is stored, what survives what, and how to back it up.


What is stored where

What Path in the container Default volume Override
Database: projects, documents, embeddings, run history /var/lib/postgresql/data contextator-pgdata CONTEXTATOR_PGDATA_VOLUME or CONTEXTATOR_PGDATA_PATH
Downloaded embedding models /app/.cache/models contextator-models CONTEXTATOR_MODELS_VOLUME / ..._PATH
Materialised sources: uploads, git checkouts, Notion pulls /data contextator-data CONTEXTATOR_DATA_VOLUME / ..._PATH
Your documentation /docs (read-only) – DOCS_HOST_PATH

What survives what

Action Data
docker compose down Kept
docker compose pull && docker compose up -d Kept
docker rm contextator, image upgrades Kept
docker compose down -v Deleted
docker volume rm contextator-pgdata Deleted

Your own documentation is never modified: /docs is mounted read-only and local sources are scanned in place.

On an installation that set DATABASE_URL to an external PostgreSQL, the first row is not yours to back up from here — the command still runs, dumps the server you named, and says so in its output; that server’s own backup regime is what actually covers it. On the -slim image, npm run backup refuses outright, because that image carries no PostgreSQL and therefore no pg_dump: run it from the default image pointed at the same DATABASE_URL instead.

Storing data in host directories

Instead of named volumes:

CONTEXTATOR_PGDATA_PATH=/srv/contextator/pgdata     # Linux: put the database on a chosen disk
CONTEXTATOR_MODELS_PATH=D:/contextator/models       # Windows (forward slashes)
CONTEXTATOR_DATA_PATH=/srv/contextator/data

Directories are created on first start and their ownership is fixed automatically. Named volumes remain the faster choice for the database on Docker Desktop; if initdb reports permission errors on a host directory, switch that one mount back to a volume.


Backing up

One command, and it takes more than the database:

docker exec contextator npm run backup -- /data/backups/contextator-$(date +%F).tar.gz
docker cp contextator:/data/backups/contextator-$(date +%F).tar.gz .

The command creates the directory it is given. Copy the resulting file off this machine — a backup sitting on the same disk as the thing it backs up is a rollback, not a backup.

What is in the archive

database.dump Every project, source, document, chunk, embedding, account, session, MCP token, audit event and logged search — including the encrypted git, Notion and Confluence tokens
data/… The files of every upload source. For uploads this is the only copy in existence; git checkouts and Notion pulls are not in here because they can be re-cloned and re-pulled
manifest.json What the archive is, what is in it, and which SECRET_KEY it needs. It is the first entry, so it can be read without unpacking the rest
README.txt The same thing in prose, for whoever opens the archive in a year

SECRET_KEY is not in the archive, on purpose

The archive records a fingerprint of SECRET_KEY — a keyed HMAC, not the key itself — never the key. A backup carrying the key would be the whole instance in one file, which is exactly what encrypting the source tokens was meant to prevent. Keep .env somewhere the archive is not, and keep a retired key around if you still hold archives taken before your last SECRET_KEY rotation and those archives carry a source’s sync credential — a private git, Notion or Confluence token — encrypted under it — see wiki/Security#rotating-secret_key for the rotation procedure and why restoring such an archive under the new key refuses.

On a schedule

There is no scheduler inside Contextator — the command is a command, and your host’s cron (or a Compose sidecar) runs it:

#!/usr/bin/env bash
set -euo pipefail
stamp=$(date +%F)
docker exec contextator npm run backup -- "/data/backups/contextator-$stamp.tar.gz"
docker cp "contextator:/data/backups/contextator-$stamp.tar.gz" /srv/backups/
docker exec contextator rm -f "/data/backups/contextator-$stamp.tar.gz"
find /srv/backups -name 'contextator-*.tar.gz' -mtime +14 -delete

set -e on the first line matters: without it, a failed backup is followed by a successful find that deletes the fortnight it was supposed to replace.


Restoring

docker exec contextator sh -c 'mkdir -p /data/backups'
docker cp contextator-2026-09-22.tar.gz contextator:/data/backups/

# Reads the manifest and every refusal, writes nothing:
docker exec contextator npm run restore -- /data/backups/contextator-2026-09-22.tar.gz --check
docker exec contextator npm run restore -- /data/backups/contextator-2026-09-22.tar.gz
docker compose restart contextator

A restore replaces what is there. It is the right thing on a fresh install and a destructive thing on a live one, and the command does not ask — --check is how you look first. Before it writes a byte it confirms the archive’s kind and version, that pg_restore is present, and that the server’s PostgreSQL major version is not older than the one the dump was taken from — any mismatch there refuses outright, with an explanation, instead of leaving a half-loaded instance. The SECRET_KEY check is narrower: it only refuses when the mismatch would lose a sync credential — a git, Notion or Confluence token nothing on this side can reissue. A webhook secret under the wrong key is not one of those; it is regenerable, so the restore counts it, reports it, and goes ahead.

A downgrade — restoring a dump from a newer PostgreSQL major into an older server — is refused outright. Going the other way, 16 to 17, is the documented upgrade path: run npm run backup, start the new major against an empty volume, then npm run restore the archive into it.

If you only want the database

The two commands underneath are unchanged and still work:

docker exec contextator pg_dump -U contextator -Fc contextator > contextator.dump
docker exec -i contextator pg_restore -U contextator -d contextator --clean --if-exists < contextator.dump

They do not carry the upload files and they do not check SECRET_KEY — which is why npm run backup and npm run restore exist.

Inspecting the database

The embedded PostgreSQL is not published, so go through the container:

docker exec -it contextator psql -U contextator
\dt                                        -- tables
SELECT name, status, document_count, chunk_count FROM projects;
SELECT name, type, status, document_count FROM document_sources;
SELECT relative_path, title, chunk_count FROM documents ORDER BY relative_path LIMIT 20;

Changing the database password

POSTGRES_PASSWORD is applied only when the cluster is first created. Later:

docker exec -it contextator psql -U contextator -c "ALTER USER contextator PASSWORD 'new-password'"

Then update .env before the next start.

Moving to another machine

  1. npm run backup and copy the archive across.
  2. Copy .env (including SECRET_KEY) and your documentation. It must be the key the archive was taken under — if you rotated SECRET_KEY after taking it, restore with the retired key instead, or take a fresh backup first.
  3. Start Contextator on the new host, then restore: mkdir -p /data/backups in the container, docker cp the archive in, npm run restore -- <file> --check, then npm run restore -- <file>, then restart.
  4. Update the URLs your agents point at — or set PUBLIC_BASE_URL and keep them stable behind a proxy.

Disk usage

  • The embedding model: ~470 MB (fp32) or ~120 MB (q8), once.
  • The database: dominated by the vectors — 384 floats per chunk plus the chunk text. A few thousand documents is typically in the hundreds of megabytes.
  • Materialised sources: as large as the uploads, checkouts and Notion pages themselves.

Arrow keys to move, Enter to open.