# Document Sources

> The six kinds of source a project can hold, the mount name every path is prefixed with, and how syncing and failures work across all of them.

- Updated: 2026-09-21
- Source: https://contextator.com/en/docs/document-sources/
- Language: en-US
- Author: Muhammet Şafak

---
A project's documentation rarely lives in one place. A **source** is one of those places, and a project
can have as many as it needs — they are merged into a single searchable endpoint.

| Type | What it is | Kept in sync by | Page |
|------|-----------|-----------------|------|
| **Local directory** | A folder mounted on the server, scanned in place. Nothing is copied | Reading it at index time | [Local Directory Source](/en/docs/local-directory-source/) |
| **Git repository** | A shallow checkout of one branch, optionally just one subdirectory | `git fetch` at the start of every run, or a push webhook | [Git Repository Source](/en/docs/git-repository-source/) |
| **Upload** | Files, folders and archives (`.zip`, `.tar.gz`, `.rar`) unpacked on the server | Nothing to sync — you upload again when it changes | [Upload Source](/en/docs/upload-source/) |
| **Notion** | Pages shared with an internal integration, rendered to Markdown | The Notion API, re-rendering only pages that changed | [Notion](/en/docs/notion/) |
| **Confluence** | Every space a Confluence Cloud or Data Center account can read, or the ones you name, rendered to Markdown | The Confluence REST API, re-rendering only pages whose version changed; Data Center can also take a webhook | [Confluence](/en/docs/confluence/) |
| **Documentation site** | A public site's own sitemap, `llms.txt` or a crawl, converted to Markdown | Re-fetching the whole site; a `sitemap.xml` with `<lastmod>` dates can be checked for changes in two requests first | [Documentation Site Source](/en/docs/documentation-site-source/) |

An **Obsidian vault** is a folder of Markdown, so it arrives as a local directory or as an upload —
with one setting flipped. See [Obsidian Vaults](/en/docs/obsidian-vaults/).

## The mount name

Every source has a **name**, and that name is the prefix of every path it contributes:

```
source "handbook"  contains  install.md   →  indexed as  handbook/install.md
source "api-repo"  contains  install.md   →  indexed as  api-repo/install.md
```

This is why two sources can hold the same file name without colliding, and why every search result tells
you which source it came from. It is also why **the name cannot be changed after creation** — every
indexed path and every path an agent has seen contains it. Everything else about a source can be edited.

Names follow the same rule as project names: lowercase letters, digits, `-` and `_`.

## Adding one

**Add source** on a project opens a dialog with a tab per kind. Common fields:

| Field | Notes |
|-------|-------|
| **Name** | The mount prefix. Immutable |
| **Label** | Free text shown in the list. Optional, changeable |
| **Content type** | `Plain Markdown / text`, `Obsidian vault`, `Notion export` or `OpenAPI / Swagger` — the last reads `.yaml`, `.yml` and `.json` as specifications and turns one into a document per operation |
| **File types** | Which extensions to index: `.md` and `.mdx` by default, optionally `.txt`, `.html`/`.htm`, `.csv`, `.docx` and `.pdf` |
| **Index now** | Queue a run as soon as the source is saved |

Git, Notion, Confluence and documentation-site sources also have a **Test connection** button that
checks credentials and reachability without indexing anything — use it before saving.

## How sources are synced

Every source is synced **at the start of every index run**, one after another, and then scanned:

1. git sources fetch the branch tip; Notion sources pull changed pages; local and upload sources have
   nothing to fetch.
2. Each source's directory is walked for the file types it selected.
3. Paths are prefixed with the source name and handed to the indexer.

Pressing **Sync** on a single source row queues the same run as **Re-index** in the header — there is
one queue per project, not per source.

## When a source fails

A source that cannot sync **reports on its own row** and the others still index. The project's status
becomes `error` with a summary:

```
2/3 sources synced; notion: API token is invalid
```

Crucially, a source whose content could not be read **at all** keeps the documents it had already
contributed. A revoked Notion share or an unreachable git host makes a source stale, never empty. Fix
the cause and press **Sync** on that row.

## Editing and removing

Changing a setting that alters what the source yields — path, branch, subdirectory, file types, content
type — queues an index run automatically, so a source is never silently stale.

Deleting a source removes its documents, its chunks and the files it materialised. It is refused while
the project is indexing.

## What gets indexed

Only the file types the source selected: `.md` and `.mdx` by default, and optionally `.txt`,
`.html`/`.htm`, `.csv`, `.docx` and `.pdf`. **Everything becomes Markdown on the way in** — the
conversion happens once, at the edge, so the chunker, the embedder and `read_document` see one format.
`.yaml`, `.yml` and `.json` are not on that list: they are readable only by the OpenAPI / Swagger
content type, and a specification it reads becomes one document per operation rather than one document.

Always skipped:

- dotfiles and dot-directories (`.git/`, `.obsidian/`, …)
- `node_modules`, `dist`, `build`, `vendor`, `__pycache__`
- symbolic links that point outside the source
- anything matching `IGNORE_GLOBS`

| Type | What it becomes | Kept | Lost |
|------|-----------------|------|------|
| `.md`, `.mdx`, `.txt` | itself, unchanged | everything | nothing |
| `.html`, `.htm` | Markdown via turndown + GFM | headings, lists, tables, code, links, `<title>` | scripts, stylesheets, `svg`, embedded image data (the `alt` text stays) |
| `.docx` | Markdown via mammoth, then the same converter | Word's own heading styles, numbered and bulleted lists, tables, links | images, footnotes, comments, tracked changes |
| `.csv` | one GFM table, `## Rows n–m` sections every 200 rows | the header above every section, quoted commas and newlines, `;`/tab/pipe delimiters | nothing of the data; cell newlines become `<br>` |
| `.pdf` | Markdown reconstructed from glyph positions | headings by font size, paragraphs rejoined across line ends and de-hyphenated, bullet and numbered lists, column-aligned tables, two-column reading order, running heads and feet dropped | footnotes, figures, and any table whose columns are not aligned |

Images and other binaries are not indexed. There is no OCR either, so a file that cannot be converted —
a scanned PDF, an encrypted one, a Word file that is all images, a damaged file of any of these types, a
page or spreadsheet that converts to no text at all — is **refused, not indexed**: the reason naming the
file is shown on its source row, and the rest of the source indexes normally. A refusal never fails the
sync or the project: on an incremental run the file keeps whatever document it already had, and a
rebuild simply leaves it out of the new generation, the same rule applied to a source that cannot be
read at all.

Conversion runs on its own worker thread, not the thread serving the dashboard and the MCP endpoint, so
one slow file never blocks a search and a parser that exhausts its heap fails only that file. Size caps
bound what a single file may cost while it converts — `MAX_CONVERTED_FILE_BYTES` for a document,
`MAX_PDF_PAGES` for a PDF that claims an implausible page count, `MAX_DOCX_UNPACKED_BYTES` against a zip
bomb inside a `.docx` — and `CONVERSION_TIMEOUT_MS` / `CONVERSION_IDLE_MS` bound how long a file, and an
idle worker, are allowed to run. Full defaults and what each one stops: [Configuration](/en/docs/configuration/#documents).

**Where a long PDF lands against `MAX_STORED_DOCUMENT_BYTES`** (1 MB of UTF-8, past which the stored
prefix is what `read_document` serves and `content_truncated` is set — the document stays fully
searchable regardless, because the cut is on the stored text and not on the chunks). Measured on
generated manuals of 80, 200, 600 and 1600 dense pages (42 lines of about 95 characters, a running head
and foot, a chapter heading every tenth page):

| Pages | Markdown | Per page | Extraction |
|-------|----------|----------|------------|
| 80 | 316 KiB | 4.0 KiB | 0.18 s |
| 200 | 795 KiB | 4.0 KiB | 0.21 s |
| 600 | 2393 KiB | 4.0 KiB | 0.61 s |
| 1600 | 6414 KiB | 4.0 KiB | 1.73 s |

The cap bites at roughly 250 dense pages — looser real-world manuals stretch that to 300–400 pages.
Extraction costs about a millisecond a page and happens once, after the content hash says the file
changed. What is capped on the way in, before any of this, is the file itself: `UPLOAD_MAX_FILE_BYTES`,
50 MB by default.
