# Indexing

> What happens during an index run, why re-indexing an unchanged project embeds nothing, how documents are chunked, and what is refused while indexing is in progress.

- Updated: 2026-09-21
- Source: https://contextator.com/en/docs/indexing/
- Language: en-US
- Author: Muhammet Şafak

---
Indexing is what turns your files into something an agent can search. It runs in the background, one
project at a time, and only does the work that is actually needed.

## What happens during a run

1. **Sync.** Every source of the project is synced in turn: git sources fetch the branch tip, Notion
   sources pull changed pages, local and upload sources have nothing to fetch. A source that fails is
   reported on its own row and the run continues with the others.
2. **Scan.** Each source's directory is walked for the file types it selected. Dotfiles,
   `node_modules`, `dist`, `build`, `vendor`, `__pycache__`, escaping symlinks and `IGNORE_GLOBS`
   matches are skipped. Every path collected is prefixed with the source name.
3. **Compare.** Every file is read and hashed (sha256). Files whose hash is unchanged are **skipped** —
   no chunking, no embedding. Changed and new files continue.
4. **Transform and chunk.** The source's content type is applied, then the file is split into chunks
   (see below).
5. **Embed.** Each chunk is embedded in batches and written to the database, replacing that document's
   previous chunks in one transaction.
6. **Clean up.** Documents whose file has disappeared are deleted — unless their source could not be
   scanned at all, in which case they are deliberately kept.
7. **Finalise.** Counts are recomputed, the project's status is set, and the run is recorded in the
   history.

## Incremental by default

A re-index of an unchanged project embeds **nothing**. The run history shows it plainly:

```
12 unchanged · 0 updated
```

This is what makes frequent re-indexing cheap — from a webhook, a cron job, or just pressing the button
whenever you are unsure.

| Action | Effect |
|--------|--------|
| **Re-index** | Incremental: only changed files are re-embedded |
| **Force re-index** | Drops every chunk of the project and rebuilds everything |

A full re-index also happens automatically when the embedding model changed — see
[Embedding Models](/en/docs/embedding-models/).

### When to force

- After changing `CHUNK_MAX_TOKENS` or `CHUNK_OVERLAP_TOKENS`, to apply them to files that did not change.
- When you suspect the index is inconsistent in a way hashing cannot see.

Changing a source's **content type** does not need a force — Contextator drops that source's hashes for
you.

## How documents are chunked

Retrieval quality is mostly chunking quality, so this part is deliberate:

- **Frontmatter is parsed**, and a `title:` in it wins over the first `# heading`, which wins over a
  prettified filename.
- **MDX is cleaned**: `import`/`export` statements, JSX comments and standalone component tags are
  removed from `.mdx` files — but never from inside fenced code.
- **The document is split at headings** `#` to `####`, and each chunk keeps a **breadcrumb** of the
  headings above it: `Guide > Install > Docker`.
- **Oversized sections are packed** from paragraphs and fenced code blocks, with a small overlap between
  consecutive chunks. Code blocks are never split mid-block unless a single block is itself too large,
  and the overlap never duplicates code.
- **Tiny fragments are merged** into their neighbour, so a lone heading never becomes a chunk of its own.
- The text handed to the embedding model is **breadcrumb + content**, so the most identifying words are
  always inside the model's window.

That breadcrumb is what you see in search results, and it is why an excerpt reads as a unit rather than
a fragment.

### Chunk size

| Setting | Default | Notes |
|---------|---------|-------|
| `CHUNK_MAX_TOKENS` | `96` | Counted with the embedding model's own tokenizer, not approximated |
| `CHUNK_OVERLAP_TOKENS` | `24` | Must be smaller than `CHUNK_MAX_TOKENS` |

The default model reads 512 tokens, so `96` is nowhere near its window — and that is deliberate.
**`96` is what measured best on the golden set**; filling the window measures *worse*, because a longer
chunk is averaged into one vector and points at nothing in particular. The budget was swept from 496 down
to 64 on the golden set, and 88 through 108 all measured the same — 96 is the middle of that plateau
rather than its edge. It is counted with the embedding model's own tokenizer rather than approximated
from the character count, because the approximation it replaced was biased by the text itself: on the
retrieval corpus it under-counted English by 17 % and Turkish by 11 %. The budget is charged for the
heading breadcrumb and, on the default model, the `passage: ` prefix it prepends to every indexed chunk
(see [How search actually works](/en/docs/embedding-models/#how-search-actually-works)) — both are part
of what the model reads, so both are counted. Raise both settings on OpenAI, whose window is 8191. Change
either, restart, then **Force re-index**; an existing project keeps its old chunks until then.

On the golden set in [`eval/`](https://github.com/Contextator/Contextator/tree/main/eval), the tokenizer
and budget work moved `recall@5` from 70.8 % to 79.2 %, and the move to `multilingual-e5-small` at 96/24
took it to 85.4 % with `recall@1` at 77.1 %.

## Watching a run

The project row shows a progress bar with the phase, files done out of total, and chunks embedded. The
phases are:

| Phase | Meaning |
|-------|---------|
| `queued` | Waiting for the indexer (which handles one project at a time) |
| `syncing` | Fetching from git / Notion, validating source roots |
| `scanning` | Walking directories and reading the existing document list |
| `embedding` | Hashing, chunking and embedding |
| `finalizing` | Recounting and writing the result |

## Run history

The **Index runs** panel keeps the last 20 runs: when, incremental or force, what changed
(`12 unchanged · 3 updated · 1 removed`), how long it took, and any error. Unlike the live progress
view, this survives restarts — it is the place to answer *"did last night's webhook actually work?"*.

Also available at `GET /api/projects/:id/runs` — see [Admin API](/en/docs/admin-api/).

## The queue

One project is indexed at a time, because embedding is CPU-bound and running two would make both
slower. A queued project shows which project it is waiting for. Queuing the same project twice returns
the run that is already in flight rather than starting a second one.

## Things that are refused during a run

To keep the index consistent, these are refused with a conflict error while a project is indexing:

- deleting the project;
- deleting one of its sources;
- committing an upload to it.

Wait for the run to finish — or let it finish and retry; the dashboard's buttons disable themselves.
