Contextator
ENTR

Operating

Indexing

What happens during an index run, why re-indexing an unchanged project embeds nothing, how documents are chunked, and what is refused while indexing is in progress.

Updated:

Indexing is what turns your files into something an agent can search. It runs in the background, one project at a time, and only does the work that is actually needed.

What happens during a run

  1. Sync. Every source of the project is synced in turn: git sources fetch the branch tip, Notion sources pull changed pages, local and upload sources have nothing to fetch. A source that fails is reported on its own row and the run continues with the others.
  2. Scan. Each source’s directory is walked for the file types it selected. Dotfiles, node_modules, dist, build, vendor, __pycache__, escaping symlinks and IGNORE_GLOBS matches are skipped. Every path collected is prefixed with the source name.
  3. Compare. Every file is read and hashed (sha256). Files whose hash is unchanged are skipped — no chunking, no embedding. Changed and new files continue.
  4. Transform and chunk. The source’s content type is applied, then the file is split into chunks (see below).
  5. Embed. Each chunk is embedded in batches and written to the database, replacing that document’s previous chunks in one transaction.
  6. Clean up. Documents whose file has disappeared are deleted — unless their source could not be scanned at all, in which case they are deliberately kept.
  7. Finalise. Counts are recomputed, the project’s status is set, and the run is recorded in the history.

Incremental by default

A re-index of an unchanged project embeds nothing. The run history shows it plainly:

12 unchanged · 0 updated

This is what makes frequent re-indexing cheap — from a webhook, a cron job, or just pressing the button whenever you are unsure.

Action Effect
Re-index Incremental: only changed files are re-embedded
Force re-index Drops every chunk of the project and rebuilds everything

A full re-index also happens automatically when the embedding model changed — see Embedding Models.

When to force

  • After changing CHUNK_MAX_TOKENS or CHUNK_OVERLAP_TOKENS, to apply them to files that did not change.
  • When you suspect the index is inconsistent in a way hashing cannot see.

Changing a source’s content type does not need a force — Contextator drops that source’s hashes for you.

How documents are chunked

Retrieval quality is mostly chunking quality, so this part is deliberate:

  • Frontmatter is parsed, and a title: in it wins over the first # heading, which wins over a prettified filename.
  • MDX is cleaned: import/export statements, JSX comments and standalone component tags are removed from .mdx files — but never from inside fenced code.
  • The document is split at headings # to ####, and each chunk keeps a breadcrumb of the headings above it: Guide > Install > Docker.
  • Oversized sections are packed from paragraphs and fenced code blocks, with a small overlap between consecutive chunks. Code blocks are never split mid-block unless a single block is itself too large, and the overlap never duplicates code.
  • Tiny fragments are merged into their neighbour, so a lone heading never becomes a chunk of its own.
  • The text handed to the embedding model is breadcrumb + content, so the most identifying words are always inside the model’s window.

That breadcrumb is what you see in search results, and it is why an excerpt reads as a unit rather than a fragment.

Chunk size

Setting Default Notes
CHUNK_MAX_TOKENS 96 Counted with the embedding model’s own tokenizer, not approximated
CHUNK_OVERLAP_TOKENS 24 Must be smaller than CHUNK_MAX_TOKENS

The default model reads 512 tokens, so 96 is nowhere near its window — and that is deliberate. 96 is what measured best on the golden set; filling the window measures worse, because a longer chunk is averaged into one vector and points at nothing in particular. The budget was swept from 496 down to 64 on the golden set, and 88 through 108 all measured the same — 96 is the middle of that plateau rather than its edge. It is counted with the embedding model’s own tokenizer rather than approximated from the character count, because the approximation it replaced was biased by the text itself: on the retrieval corpus it under-counted English by 17 % and Turkish by 11 %. The budget is charged for the heading breadcrumb and, on the default model, the passage: prefix it prepends to every indexed chunk (see How search actually works) — both are part of what the model reads, so both are counted. Raise both settings on OpenAI, whose window is 8191. Change either, restart, then Force re-index; an existing project keeps its old chunks until then.

On the golden set in eval/, the tokenizer and budget work moved recall@5 from 70.8 % to 79.2 %, and the move to multilingual-e5-small at 96/24 took it to 85.4 % with recall@1 at 77.1 %.

Watching a run

The project row shows a progress bar with the phase, files done out of total, and chunks embedded. The phases are:

Phase Meaning
queued Waiting for the indexer (which handles one project at a time)
syncing Fetching from git / Notion, validating source roots
scanning Walking directories and reading the existing document list
embedding Hashing, chunking and embedding
finalizing Recounting and writing the result

Run history

The Index runs panel keeps the last 20 runs: when, incremental or force, what changed (12 unchanged · 3 updated · 1 removed), how long it took, and any error. Unlike the live progress view, this survives restarts — it is the place to answer “did last night’s webhook actually work?”.

Also available at GET /api/projects/:id/runs — see Admin API.

The queue

One project is indexed at a time, because embedding is CPU-bound and running two would make both slower. A queued project shows which project it is waiting for. Queuing the same project twice returns the run that is already in flight rather than starting a second one.

Things that are refused during a run

To keep the index consistent, these are refused with a conflict error while a project is indexing:

  • deleting the project;
  • deleting one of its sources;
  • committing an upload to it.

Wait for the run to finish — or let it finish and retry; the dashboard’s buttons disable themselves.

Arrow keys to move, Enter to open.