# Documentation Site Source

> Indexing a public site's own sitemap, llms.txt or a crawl — detection, the five crawl ceilings, robots.txt, and how freshness is checked.

- Updated: 2026-09-25
- Source: https://contextator.com/en/docs/documentation-site-source/
- Language: en-US
- Author: Muhammet Şafak

---
Index a public documentation site by its own sitemap, its own `llms.txt`, or by crawling it from one
entry point — no login, no credential, nothing installed on the target.

## Adding one

**Add source → Documentation site**

| Field | Notes |
|-------|-------|
| **Name** | The mount prefix. Immutable |
| **Entry point** | A `sitemap.xml`, an `llms.txt`, or an ordinary page to crawl from |
| **Entry format** | *Detect from the document* by default. Set it by hand — sitemap, llms.txt or crawl — when detection refuses |
| **Index now** | Queue a run as soon as the source is saved |

Press **Test connection** to fetch the entry point, detect which of the three formats it is, and report
what was found — a page count from a sitemap, a link count from `llms.txt`, or, for a crawl, the number
of links the start page itself offers.

## What the entry point can be

Contextator reads the document at the URL you gave it, not its extension, so any of the three works at
any path:

- **`sitemap.xml`** — every `<loc>` is a page to fetch. A `<sitemapindex>` that points at other sitemaps
  is followed automatically, as deep as `WEB_MAX_DEPTH` and at most `WEB_MAX_PAGES` nested sitemaps read.
- **`llms.txt`** — the emerging convention for an LLM-readable index of a site. Every Markdown link on
  the page is a page to fetch.
- **A single page** — anything else starts a **crawl**: breadth-first from that URL, following links
  that share its host and stay under its path.

Detection reads the document, not the URL. When the document at the entry point matches none of the
three — a sitemap that redirects to a landing page, an `llms.txt` served as HTML — the source refuses
and names the **Entry format** field rather than guessing, because guessing produces a source that
indexes one page, or none, and reports success. Set **Entry format** by hand in that case: it is a
working override, not just a display of what was detected.

## The five ceilings

A crawl is bounded by five limits, all instance-wide environment variables — no per-source override:

| Ceiling | Variable | Default |
|---------|----------|---------|
| Pages per source | `WEB_MAX_PAGES` | 1000 |
| Link depth from the entry point | `WEB_MAX_DEPTH` | 10 |
| Delay between requests | `WEB_REQUEST_DELAY_MS` | 500 |
| Total crawl time budget | `WEB_CRAWL_BUDGET_MS` | 900000 (15 minutes) |
| Respect `robots.txt` | `WEB_RESPECT_ROBOTS` | 1 (on) |

Hitting a ceiling ends the crawl where it stands rather than failing the sync — the source indexes
whatever pages it reached and the run says which limit stopped it. All five are instance-wide `WEB_*`
variables in `.env`, listed in the shipped
[`.env.example`](https://github.com/Contextator/Contextator/blob/main/.env.example) — narrowing the
entry point, or splitting the site across several sources, is the better fix for a ceiling that was
reached before raising one of them.

**`WEB_MAX_PAGES` counts pages fetched, not pages indexed**, deliberately: a page refused for having no
text outside its scripts, a `404` from a stale sitemap link, or a failed connection was still served by
somebody's web server, so it counts against the one setting that exists to protect a host nobody here has
an account with. A URL that is never requested at all is never charged — one `robots.txt` disallows, or
one that sits on another host, is a decision taken locally before anything leaves the process.

## `robots.txt`

Obeyed by default (`WEB_RESPECT_ROBOTS=1`):

- A `Crawl-delay` directive only ever **raises** the pacing above `WEB_REQUEST_DELAY_MS` — it cannot
  lower it.
- `robots.txt` answering `404` is read as **no restriction** — the crawl proceeds.
- `robots.txt` failing to load at all — a `5xx` or a connection error — **fails the sync**, rather than
  crawling a site whose rules could not be read.

There is no headless browser: a page whose content is rendered by client-side JavaScript rather than
present in the HTML response is refused by name, the same way an unrenderable file is on any other
source.

## What a page becomes

Every fetched page goes through the same HTML→Markdown transform as an `.html` file from any other
source — see [Content Types](/en/docs/content-types/). The URL becomes the document's path: the content
type of the response decides the extension, not whatever suffix the URL happened to end in, and a query
string is hashed into the path rather than dropped, so two URLs that differ only by query string do not
collide.

## How it stays fresh

Only a **`sitemap.xml`** source gets a cheap freshness check: two requests, `robots.txt` then the
sitemap itself, comparing entry count and the newest `<lastmod>` date against the last run. If neither
moved, nothing is re-fetched. A sitemap that carries no `<lastmod>` at all has nothing to compare, so it
is re-fetched in full every run, the same as the other two shapes below.

**`llms.txt`** and **crawl** sources have no equivalent signal to check cheaply, so every scheduled run
re-fetches the source in full, within the same five ceilings. See [Indexing](/en/docs/indexing/) for how
the schedule itself is set.

## Common problems

| Symptom | Cause and fix |
|---------|---------------|
| The source refuses and names **Entry format** | Detection could not tell what the entry point is — a sitemap that redirects to a landing page, an `llms.txt` served as HTML. Set the format by hand |
| A page is missing from the index | It sits past `WEB_MAX_PAGES` or `WEB_MAX_DEPTH`, was excluded by `robots.txt`, or renders its content with client-side JavaScript |
| The run stopped at a ceiling | The source's row names which one — pages, depth or the time budget. Narrow the entry point or split the site across sources before raising the setting |
| The sync failed on `robots.txt` | It could not be read at all (`5xx` or a connection error), which is not treated as permission. A `404` would have been fine |
| A page is refused as *served as …, which is not a page this source can read* | The host sends a content type this source does not read — `application/octet-stream` on a `.md` file is the common one. Only HTML, Markdown and plain text are read |
| The sync fails and says it is refusing to delete existing documents | The run indexed nothing while the source already held documents — a sitemap a CDN briefly answered with a landing page, a renamed `llms.txt`. The documents are kept rather than deleted as gone; check the entry point, then sync again |
