Indexing is what turns your files into something an agent can search. It runs in the background, one project at a time, and only does the work that is actually needed.
What happens during a run
- Sync. Every source of the project is synced in turn: git sources fetch the branch tip, Notion sources pull changed pages, local and upload sources have nothing to fetch. A source that fails is reported on its own row and the run continues with the others.
- Scan. Each source’s directory is walked for the file types it selected. Dotfiles,
node_modules,dist,build,vendor,__pycache__, escaping symlinks andIGNORE_GLOBSmatches are skipped. Every path collected is prefixed with the source name. - Compare. Every file is read and hashed (sha256). Files whose hash is unchanged are skipped — no chunking, no embedding. Changed and new files continue.
- Transform and chunk. The source’s content type is applied, then the file is split into chunks (see below).
- Embed. Each chunk is embedded in batches and written to the database, replacing that document’s previous chunks in one transaction.
- Clean up. Documents whose file has disappeared are deleted — unless their source could not be scanned at all, in which case they are deliberately kept.
- Finalise. Counts are recomputed, the project’s status is set, and the run is recorded in the history.
Incremental by default
A re-index of an unchanged project embeds nothing. The run history shows it plainly:
12 unchanged · 0 updated
This is what makes frequent re-indexing cheap — from a webhook, a cron job, or just pressing the button whenever you are unsure.
| Action | Effect |
|---|---|
| Re-index | Incremental: only changed files are re-embedded |
| Force re-index | Drops every chunk of the project and rebuilds everything |
A full re-index also happens automatically when the embedding model changed — see Embedding Models.
When to force
- After changing
CHUNK_MAX_TOKENSorCHUNK_OVERLAP_TOKENS, to apply them to files that did not change. - When you suspect the index is inconsistent in a way hashing cannot see.
Changing a source’s content type does not need a force — Contextator drops that source’s hashes for you.
How documents are chunked
Retrieval quality is mostly chunking quality, so this part is deliberate:
- Frontmatter is parsed, and a
title:in it wins over the first# heading, which wins over a prettified filename. - MDX is cleaned:
import/exportstatements, JSX comments and standalone component tags are removed from.mdxfiles — but never from inside fenced code. - The document is split at headings
#to####, and each chunk keeps a breadcrumb of the headings above it:Guide > Install > Docker. - Oversized sections are packed from paragraphs and fenced code blocks, with a small overlap between consecutive chunks. Code blocks are never split mid-block unless a single block is itself too large, and the overlap never duplicates code.
- Tiny fragments are merged into their neighbour, so a lone heading never becomes a chunk of its own.
- The text handed to the embedding model is breadcrumb + content, so the most identifying words are always inside the model’s window.
That breadcrumb is what you see in search results, and it is why an excerpt reads as a unit rather than a fragment.
Chunk size
| Setting | Default | Notes |
|---|---|---|
CHUNK_MAX_TOKENS |
96 |
Counted with the embedding model’s own tokenizer, not approximated |
CHUNK_OVERLAP_TOKENS |
24 |
Must be smaller than CHUNK_MAX_TOKENS |
The default model reads 512 tokens, so 96 is nowhere near its window — and that is deliberate.
96 is what measured best on the golden set; filling the window measures worse, because a longer
chunk is averaged into one vector and points at nothing in particular. The budget was swept from 496 down
to 64 on the golden set, and 88 through 108 all measured the same — 96 is the middle of that plateau
rather than its edge. It is counted with the embedding model’s own tokenizer rather than approximated
from the character count, because the approximation it replaced was biased by the text itself: on the
retrieval corpus it under-counted English by 17 % and Turkish by 11 %. The budget is charged for the
heading breadcrumb and, on the default model, the passage: prefix it prepends to every indexed chunk
(see How search actually works) — both are part
of what the model reads, so both are counted. Raise both settings on OpenAI, whose window is 8191. Change
either, restart, then Force re-index; an existing project keeps its old chunks until then.
On the golden set in eval/, the tokenizer
and budget work moved recall@5 from 70.8 % to 79.2 %, and the move to multilingual-e5-small at 96/24
took it to 85.4 % with recall@1 at 77.1 %.
Watching a run
The project row shows a progress bar with the phase, files done out of total, and chunks embedded. The phases are:
| Phase | Meaning |
|---|---|
queued |
Waiting for the indexer (which handles one project at a time) |
syncing |
Fetching from git / Notion, validating source roots |
scanning |
Walking directories and reading the existing document list |
embedding |
Hashing, chunking and embedding |
finalizing |
Recounting and writing the result |
Run history
The Index runs panel keeps the last 20 runs: when, incremental or force, what changed
(12 unchanged · 3 updated · 1 removed), how long it took, and any error. Unlike the live progress
view, this survives restarts — it is the place to answer “did last night’s webhook actually work?”.
Also available at GET /api/projects/:id/runs — see Admin API.
The queue
One project is indexed at a time, because embedding is CPU-bound and running two would make both slower. A queued project shows which project it is waiting for. Queuing the same project twice returns the run that is already in flight rather than starting a second one.
Things that are refused during a run
To keep the index consistent, these are refused with a conflict error while a project is indexing:
- deleting the project;
- deleting one of its sources;
- committing an upload to it.
Wait for the run to finish — or let it finish and retry; the dashboard’s buttons disable themselves.