Index a public documentation site by its own sitemap, its own llms.txt, or by crawling it from one
entry point — no login, no credential, nothing installed on the target.
Adding one
Add source → Documentation site
| Field | Notes |
|---|---|
| Name | The mount prefix. Immutable |
| Entry point | A sitemap.xml, an llms.txt, or an ordinary page to crawl from |
| Entry format | Detect from the document by default. Set it by hand — sitemap, llms.txt or crawl — when detection refuses |
| Index now | Queue a run as soon as the source is saved |
Press Test connection to fetch the entry point, detect which of the three formats it is, and report
what was found — a page count from a sitemap, a link count from llms.txt, or, for a crawl, the number
of links the start page itself offers.
What the entry point can be
Contextator reads the document at the URL you gave it, not its extension, so any of the three works at any path:
sitemap.xml— every<loc>is a page to fetch. A<sitemapindex>that points at other sitemaps is followed automatically, as deep asWEB_MAX_DEPTHand at mostWEB_MAX_PAGESnested sitemaps read.llms.txt— the emerging convention for an LLM-readable index of a site. Every Markdown link on the page is a page to fetch.- A single page — anything else starts a crawl: breadth-first from that URL, following links that share its host and stay under its path.
Detection reads the document, not the URL. When the document at the entry point matches none of the
three — a sitemap that redirects to a landing page, an llms.txt served as HTML — the source refuses
and names the Entry format field rather than guessing, because guessing produces a source that
indexes one page, or none, and reports success. Set Entry format by hand in that case: it is a
working override, not just a display of what was detected.
The five ceilings
A crawl is bounded by five limits, all instance-wide environment variables — no per-source override:
| Ceiling | Variable | Default |
|---|---|---|
| Pages per source | WEB_MAX_PAGES |
1000 |
| Link depth from the entry point | WEB_MAX_DEPTH |
10 |
| Delay between requests | WEB_REQUEST_DELAY_MS |
500 |
| Total crawl time budget | WEB_CRAWL_BUDGET_MS |
900000 (15 minutes) |
Respect robots.txt |
WEB_RESPECT_ROBOTS |
1 (on) |
Hitting a ceiling ends the crawl where it stands rather than failing the sync — the source indexes
whatever pages it reached and the run says which limit stopped it. All five are instance-wide WEB_*
variables in .env, listed in the shipped
.env.example — narrowing the
entry point, or splitting the site across several sources, is the better fix for a ceiling that was
reached before raising one of them.
WEB_MAX_PAGES counts pages fetched, not pages indexed, deliberately: a page refused for having no
text outside its scripts, a 404 from a stale sitemap link, or a failed connection was still served by
somebody’s web server, so it counts against the one setting that exists to protect a host nobody here has
an account with. A URL that is never requested at all is never charged — one robots.txt disallows, or
one that sits on another host, is a decision taken locally before anything leaves the process.
robots.txt
Obeyed by default (WEB_RESPECT_ROBOTS=1):
- A
Crawl-delaydirective only ever raises the pacing aboveWEB_REQUEST_DELAY_MS— it cannot lower it. robots.txtanswering404is read as no restriction — the crawl proceeds.robots.txtfailing to load at all — a5xxor a connection error — fails the sync, rather than crawling a site whose rules could not be read.
There is no headless browser: a page whose content is rendered by client-side JavaScript rather than present in the HTML response is refused by name, the same way an unrenderable file is on any other source.
What a page becomes
Every fetched page goes through the same HTML→Markdown transform as an .html file from any other
source — see Content Types. The URL becomes the document’s path: the content
type of the response decides the extension, not whatever suffix the URL happened to end in, and a query
string is hashed into the path rather than dropped, so two URLs that differ only by query string do not
collide.
How it stays fresh
Only a sitemap.xml source gets a cheap freshness check: two requests, robots.txt then the
sitemap itself, comparing entry count and the newest <lastmod> date against the last run. If neither
moved, nothing is re-fetched. A sitemap that carries no <lastmod> at all has nothing to compare, so it
is re-fetched in full every run, the same as the other two shapes below.
llms.txt and crawl sources have no equivalent signal to check cheaply, so every scheduled run
re-fetches the source in full, within the same five ceilings. See Indexing for how
the schedule itself is set.
Common problems
| Symptom | Cause and fix |
|---|---|
| The source refuses and names Entry format | Detection could not tell what the entry point is — a sitemap that redirects to a landing page, an llms.txt served as HTML. Set the format by hand |
| A page is missing from the index | It sits past WEB_MAX_PAGES or WEB_MAX_DEPTH, was excluded by robots.txt, or renders its content with client-side JavaScript |
| The run stopped at a ceiling | The source’s row names which one — pages, depth or the time budget. Narrow the entry point or split the site across sources before raising the setting |
The sync failed on robots.txt |
It could not be read at all (5xx or a connection error), which is not treated as permission. A 404 would have been fine |
| A page is refused as served as …, which is not a page this source can read | The host sends a content type this source does not read — application/octet-stream on a .md file is the common one. Only HTML, Markdown and plain text are read |
| The sync fails and says it is refusing to delete existing documents | The run indexed nothing while the source already held documents — a sitemap a CDN briefly answered with a landing page, a renamed llms.txt. The documents are kept rather than deleted as gone; check the entry point, then sync again |