Contextator
ENTR

Sources

Document Sources

The six kinds of source a project can hold, the mount name every path is prefixed with, and how syncing and failures work across all of them.

Updated:

A project’s documentation rarely lives in one place. A source is one of those places, and a project can have as many as it needs — they are merged into a single searchable endpoint.

Type What it is Kept in sync by Page
Local directory A folder mounted on the server, scanned in place. Nothing is copied Reading it at index time Local Directory Source
Git repository A shallow checkout of one branch, optionally just one subdirectory git fetch at the start of every run, or a push webhook Git Repository Source
Upload Files, folders and archives (.zip, .tar.gz, .rar) unpacked on the server Nothing to sync — you upload again when it changes Upload Source
Notion Pages shared with an internal integration, rendered to Markdown The Notion API, re-rendering only pages that changed Notion
Confluence Every space a Confluence Cloud or Data Center account can read, or the ones you name, rendered to Markdown The Confluence REST API, re-rendering only pages whose version changed; Data Center can also take a webhook Confluence
Documentation site A public site’s own sitemap, llms.txt or a crawl, converted to Markdown Re-fetching the whole site; a sitemap.xml with <lastmod> dates can be checked for changes in two requests first Documentation Site Source

An Obsidian vault is a folder of Markdown, so it arrives as a local directory or as an upload — with one setting flipped. See Obsidian Vaults.

The mount name

Every source has a name, and that name is the prefix of every path it contributes:

source "handbook"  contains  install.md   →  indexed as  handbook/install.md
source "api-repo"  contains  install.md   →  indexed as  api-repo/install.md

This is why two sources can hold the same file name without colliding, and why every search result tells you which source it came from. It is also why the name cannot be changed after creation — every indexed path and every path an agent has seen contains it. Everything else about a source can be edited.

Names follow the same rule as project names: lowercase letters, digits, - and _.

Adding one

Add source on a project opens a dialog with a tab per kind. Common fields:

Field Notes
Name The mount prefix. Immutable
Label Free text shown in the list. Optional, changeable
Content type Plain Markdown / text, Obsidian vault, Notion export or OpenAPI / Swagger — the last reads .yaml, .yml and .json as specifications and turns one into a document per operation
File types Which extensions to index: .md and .mdx by default, optionally .txt, .html/.htm, .csv, .docx and .pdf
Index now Queue a run as soon as the source is saved

Git, Notion, Confluence and documentation-site sources also have a Test connection button that checks credentials and reachability without indexing anything — use it before saving.

How sources are synced

Every source is synced at the start of every index run, one after another, and then scanned:

  1. git sources fetch the branch tip; Notion sources pull changed pages; local and upload sources have nothing to fetch.
  2. Each source’s directory is walked for the file types it selected.
  3. Paths are prefixed with the source name and handed to the indexer.

Pressing Sync on a single source row queues the same run as Re-index in the header — there is one queue per project, not per source.

When a source fails

A source that cannot sync reports on its own row and the others still index. The project’s status becomes error with a summary:

2/3 sources synced; notion: API token is invalid

Crucially, a source whose content could not be read at all keeps the documents it had already contributed. A revoked Notion share or an unreachable git host makes a source stale, never empty. Fix the cause and press Sync on that row.

Editing and removing

Changing a setting that alters what the source yields — path, branch, subdirectory, file types, content type — queues an index run automatically, so a source is never silently stale.

Deleting a source removes its documents, its chunks and the files it materialised. It is refused while the project is indexing.

What gets indexed

Only the file types the source selected: .md and .mdx by default, and optionally .txt, .html/.htm, .csv, .docx and .pdf. Everything becomes Markdown on the way in — the conversion happens once, at the edge, so the chunker, the embedder and read_document see one format. .yaml, .yml and .json are not on that list: they are readable only by the OpenAPI / Swagger content type, and a specification it reads becomes one document per operation rather than one document.

Always skipped:

  • dotfiles and dot-directories (.git/, .obsidian/, …)
  • node_modules, dist, build, vendor, __pycache__
  • symbolic links that point outside the source
  • anything matching IGNORE_GLOBS
Type What it becomes Kept Lost
.md, .mdx, .txt itself, unchanged everything nothing
.html, .htm Markdown via turndown + GFM headings, lists, tables, code, links, <title> scripts, stylesheets, svg, embedded image data (the alt text stays)
.docx Markdown via mammoth, then the same converter Word’s own heading styles, numbered and bulleted lists, tables, links images, footnotes, comments, tracked changes
.csv one GFM table, ## Rows n–m sections every 200 rows the header above every section, quoted commas and newlines, ;/tab/pipe delimiters nothing of the data; cell newlines become <br>
.pdf Markdown reconstructed from glyph positions headings by font size, paragraphs rejoined across line ends and de-hyphenated, bullet and numbered lists, column-aligned tables, two-column reading order, running heads and feet dropped footnotes, figures, and any table whose columns are not aligned

Images and other binaries are not indexed. There is no OCR either, so a file that cannot be converted — a scanned PDF, an encrypted one, a Word file that is all images, a damaged file of any of these types, a page or spreadsheet that converts to no text at all — is refused, not indexed: the reason naming the file is shown on its source row, and the rest of the source indexes normally. A refusal never fails the sync or the project: on an incremental run the file keeps whatever document it already had, and a rebuild simply leaves it out of the new generation, the same rule applied to a source that cannot be read at all.

Conversion runs on its own worker thread, not the thread serving the dashboard and the MCP endpoint, so one slow file never blocks a search and a parser that exhausts its heap fails only that file. Size caps bound what a single file may cost while it converts — MAX_CONVERTED_FILE_BYTES for a document, MAX_PDF_PAGES for a PDF that claims an implausible page count, MAX_DOCX_UNPACKED_BYTES against a zip bomb inside a .docx — and CONVERSION_TIMEOUT_MS / CONVERSION_IDLE_MS bound how long a file, and an idle worker, are allowed to run. Full defaults and what each one stops: Configuration.

Where a long PDF lands against MAX_STORED_DOCUMENT_BYTES (1 MB of UTF-8, past which the stored prefix is what read_document serves and content_truncated is set — the document stays fully searchable regardless, because the cut is on the stored text and not on the chunks). Measured on generated manuals of 80, 200, 600 and 1600 dense pages (42 lines of about 95 characters, a running head and foot, a chapter heading every tenth page):

Pages Markdown Per page Extraction
80 316 KiB 4.0 KiB 0.18 s
200 795 KiB 4.0 KiB 0.21 s
600 2393 KiB 4.0 KiB 0.61 s
1600 6414 KiB 4.0 KiB 1.73 s

The cap bites at roughly 250 dense pages — looser real-world manuals stretch that to 300–400 pages. Extraction costs about a millisecond a page and happens once, after the content hash says the file changed. What is capped on the way in, before any of this, is the file itself: UPLOAD_MAX_FILE_BYTES, 50 MB by default.

Arrow keys to move, Enter to open.