Contextator
ENTR

Sources

Documentation Site Source

Indexing a public site's own sitemap, llms.txt or a crawl — detection, the five crawl ceilings, robots.txt, and how freshness is checked.

Updated:

Index a public documentation site by its own sitemap, its own llms.txt, or by crawling it from one entry point — no login, no credential, nothing installed on the target.

Adding one

Add source → Documentation site

Field Notes
Name The mount prefix. Immutable
Entry point A sitemap.xml, an llms.txt, or an ordinary page to crawl from
Entry format Detect from the document by default. Set it by hand — sitemap, llms.txt or crawl — when detection refuses
Index now Queue a run as soon as the source is saved

Press Test connection to fetch the entry point, detect which of the three formats it is, and report what was found — a page count from a sitemap, a link count from llms.txt, or, for a crawl, the number of links the start page itself offers.

What the entry point can be

Contextator reads the document at the URL you gave it, not its extension, so any of the three works at any path:

  • sitemap.xml — every <loc> is a page to fetch. A <sitemapindex> that points at other sitemaps is followed automatically, as deep as WEB_MAX_DEPTH and at most WEB_MAX_PAGES nested sitemaps read.
  • llms.txt — the emerging convention for an LLM-readable index of a site. Every Markdown link on the page is a page to fetch.
  • A single page — anything else starts a crawl: breadth-first from that URL, following links that share its host and stay under its path.

Detection reads the document, not the URL. When the document at the entry point matches none of the three — a sitemap that redirects to a landing page, an llms.txt served as HTML — the source refuses and names the Entry format field rather than guessing, because guessing produces a source that indexes one page, or none, and reports success. Set Entry format by hand in that case: it is a working override, not just a display of what was detected.

The five ceilings

A crawl is bounded by five limits, all instance-wide environment variables — no per-source override:

Ceiling Variable Default
Pages per source WEB_MAX_PAGES 1000
Link depth from the entry point WEB_MAX_DEPTH 10
Delay between requests WEB_REQUEST_DELAY_MS 500
Total crawl time budget WEB_CRAWL_BUDGET_MS 900000 (15 minutes)
Respect robots.txt WEB_RESPECT_ROBOTS 1 (on)

Hitting a ceiling ends the crawl where it stands rather than failing the sync — the source indexes whatever pages it reached and the run says which limit stopped it. All five are instance-wide WEB_* variables in .env, listed in the shipped .env.example — narrowing the entry point, or splitting the site across several sources, is the better fix for a ceiling that was reached before raising one of them.

WEB_MAX_PAGES counts pages fetched, not pages indexed, deliberately: a page refused for having no text outside its scripts, a 404 from a stale sitemap link, or a failed connection was still served by somebody’s web server, so it counts against the one setting that exists to protect a host nobody here has an account with. A URL that is never requested at all is never charged — one robots.txt disallows, or one that sits on another host, is a decision taken locally before anything leaves the process.

robots.txt

Obeyed by default (WEB_RESPECT_ROBOTS=1):

  • A Crawl-delay directive only ever raises the pacing above WEB_REQUEST_DELAY_MS — it cannot lower it.
  • robots.txt answering 404 is read as no restriction — the crawl proceeds.
  • robots.txt failing to load at all — a 5xx or a connection error — fails the sync, rather than crawling a site whose rules could not be read.

There is no headless browser: a page whose content is rendered by client-side JavaScript rather than present in the HTML response is refused by name, the same way an unrenderable file is on any other source.

What a page becomes

Every fetched page goes through the same HTML→Markdown transform as an .html file from any other source — see Content Types. The URL becomes the document’s path: the content type of the response decides the extension, not whatever suffix the URL happened to end in, and a query string is hashed into the path rather than dropped, so two URLs that differ only by query string do not collide.

How it stays fresh

Only a sitemap.xml source gets a cheap freshness check: two requests, robots.txt then the sitemap itself, comparing entry count and the newest <lastmod> date against the last run. If neither moved, nothing is re-fetched. A sitemap that carries no <lastmod> at all has nothing to compare, so it is re-fetched in full every run, the same as the other two shapes below.

llms.txt and crawl sources have no equivalent signal to check cheaply, so every scheduled run re-fetches the source in full, within the same five ceilings. See Indexing for how the schedule itself is set.

Common problems

Symptom Cause and fix
The source refuses and names Entry format Detection could not tell what the entry point is — a sitemap that redirects to a landing page, an llms.txt served as HTML. Set the format by hand
A page is missing from the index It sits past WEB_MAX_PAGES or WEB_MAX_DEPTH, was excluded by robots.txt, or renders its content with client-side JavaScript
The run stopped at a ceiling The source’s row names which one — pages, depth or the time budget. Narrow the entry point or split the site across sources before raising the setting
The sync failed on robots.txt It could not be read at all (5xx or a connection error), which is not treated as permission. A 404 would have been fine
A page is refused as served as …, which is not a page this source can read The host sends a content type this source does not read — application/octet-stream on a .md file is the common one. Only HTML, Markdown and plain text are read
The sync fails and says it is refusing to delete existing documents The run indexed nothing while the source already held documents — a sitemap a CDN briefly answered with a landing page, a renamed llms.txt. The documents are kept rather than deleted as gone; check the entry point, then sync again

Arrow keys to move, Enter to open.