# MCP Tools

> The three read-only tools every project publishes — search_docs, list_topics and read_document — their arguments, their limits, and what they cannot do.

- Updated: 2026-09-21
- Source: https://contextator.com/en/docs/mcp-tools/
- Language: en-US
- Author: Muhammet Şafak

---
Every project publishes the same three tools, scoped to its own documents. They are read-only: an agent
can search and read, never write.

When a client connects, the server also sends **instructions** naming the project, how many documents and
chunks it holds, and when to use which tool — which is why you can simply ask a question instead of
telling the agent which tool to call.

---

## `search_docs`

Hybrid search over the project's chunks: meaning and exact wording are searched at once and fused by
rank, so `HALYARD_DISPATCH_TIMEOUT` finds its page as readily as a question does.

| Argument | Type | Default |
|----------|------|---------|
| `query` | string (1–2000 characters) | required |
| `limit` | integer 1–20 | 5 |
| `source` | string — one source name, the first segment of the paths `list_topics` shows | every source |
| `path_prefix` | string — e.g. `handbook/operations` | the whole project |
| `version` | string — an exact release label as the project set it, e.g. `v3` | every version |

The last three narrow the search and are all optional; omitting them searches everything, as it always
did. An unknown `source` or `version` is answered with the ones this project actually has, never with an
empty page — there is no ordering and no `latest`.

Returns ranked excerpts, each with its file path, heading breadcrumb and similarity score:

```
Found 3 results for "how do I rotate the signing key" in project "handbook":

### 1. runbooks/keys.md — Runbooks > Signing keys > Rotation (score 0.871)
Rotate the signing key by generating a new pair with `keytool`, publishing the public
half to the JWKS endpoint, and keeping the previous key active for 24 hours…

### 2. security/policy.md — Security > Key material (score 0.842)
…
```

Because the breadcrumb is part of what gets embedded, a question phrased like a heading tends to find
the right section directly.

Each excerpt is rendered with the chunk before and after it, marked with a leading and trailing `…`, so
an agent usually does not have to spend a `read_document` call to see the sentence a chunk boundary cut
in half. At most two excerpts come from any one document, because five results that are five consecutive
chunks of one page answer the question once and crowd out four other pages. The whole answer is capped at
`SEARCH_MAX_RESULT_CHARS` and says so when it cuts — see [Configuration](/en/docs/configuration/).

**Special answers instead of results:**

- The project has nothing indexed yet → a message saying so, and to index it.
- The project was indexed with a different embedding model → an error asking you to re-index, rather
  than confidently wrong matches.
- Nothing clears the relevance floor, a cosine similarity of `0.82` by default → *no good match* and a
  pointer at `list_topics`, rather than the least bad hit it found. See
  [Embedding Models](/en/docs/embedding-models/) for what the numbers mean.

## `list_topics`

| Argument | Type | Default |
|----------|------|---------|
| `cursor` | string — the `next_cursor` a previous call ended with | start of the listing |
| `limit` | integer 1–1000 | 200 |

One call returns `limit` documents (200 unless said otherwise, at most 1000). Returns the page grouped
by directory:

```
Project "handbook": 84 documents, 412 chunks
Sources (the first path segment): policies (local), api (git: API reference), notion (notion)

policies/ — 12 documents, 61 chunks
  • policies/onboarding.md — Onboarding (9 chunks)
  • policies/travel.md — Travel policy (4 chunks)

api/reference/ — 31 documents, 180 chunks
  • api/reference/auth.md — Authentication (12 chunks)
  …
```

Useful to an agent that wants to know what exists before searching, and to you when you want to check
what was actually indexed. When more documents remain past the page, the answer ends with a
`next_cursor` value and a line telling the agent to call `list_topics` again with that value as
`cursor`.

## `read_document`

| Argument | Type | Default |
|----------|------|---------|
| `path` | string — exactly as shown by `search_docs` or `list_topics` | required |
| `heading` | string — a heading breadcrumb out of a search result, e.g. `Guide > Install > Docker` | the whole document |
| `from`, `to` | integer — a chunk range, 0-based and inclusive | the whole document |
| `max_tokens` | integer 200–20000 | 4000 |

Returns the Markdown of one indexed file:

```
File: api/reference/auth.md
Title: Authentication

---

# Authentication
…
```

Pass `heading` with a breadcrumb from a search result to read only that section and the subsections
under it, or `from`/`to` for a chunk range — either is far cheaper than a whole page.

Only paths that were **indexed for that project** are served — never an arbitrary filesystem path. Output
is capped at `max_tokens`, counted with the embedding model's own tokenizer, and says where it cut and
how to ask for the rest.

**Reading a document is not the same as reading a file.** `read_document` serves the text this server
indexed, stored beside the chunks, and never touches the filesystem. So a document stays readable after
its file is renamed, its git checkout is re-cloned, or the whole source directory is unmounted, and a
page that was *never* indexed can never be served by mistake. The text is stored as the indexer
transformed it, which is what keeps `read_document` and `search_docs` from ever disagreeing about what a
page says; the price is that an Obsidian note's original `[[wikilink]]` syntax is not readable through
MCP, only the Markdown link it became. Documents indexed by an older version hold no stored text and are
read from disk until their next index run, which is the one case where a missing file is still an error.

---

## Paths and sources

Every path begins with the name of the source it came from:

```
handbook/install.md      ← source "handbook"
api-repo/install.md      ← source "api-repo"
```

That is what tells you where an answer came from, and what lets two sources hold the same file name.
See [Document Sources](/en/docs/document-sources/).

## Getting better answers

- **Ask questions, but identifiers work too.** Search is hybrid, so an environment variable, a header
  or an error code finds its page by exact wording; for everything else a full question carries more
  signal for the meaning half than two words do.
- **Ask in the language the page is written in.** The default model covers 100 languages including
  Turkish, but retrieval does not cross between them — that is a measured limit of the model, not a
  switch.
- **Narrow it when you know where the answer lives.** `source` and `path_prefix` cut the search to one
  source or one directory, and `version` to one release, so `v2` and `v3` of the same documentation do
  not answer for each other.
- **Raise `limit`** when a topic is spread over many files — an agent can ask for up to 20 excerpts.
- **Structure your documentation with headings.** Chunks are cut at headings and carry their breadcrumb,
  so good headings directly improve retrieval.
- **Keep unrelated bodies of knowledge in separate projects.** Precision drops when one endpoint mixes
  a handbook, an API reference and a journal.
- **Tell the agent to cite the file path.** The server already asks it to; it is worth reinforcing in
  your own prompt.

## Structured output

Set `MCP_STRUCTURED_OUTPUT=1` and every tool also publishes an `outputSchema` and returns
`structuredContent` alongside the same text (MCP 2025-06-18). It is **off by default**: a client that
reads structured content can stop reading the text, and Claude Code does exactly that once a result
carries `structuredContent`
([anthropics/claude-code#55677](https://github.com/anthropics/claude-code/issues/55677),
[#79944](https://github.com/anthropics/claude-code/issues/79944)), so the prose guiding the model on how
to read excerpts would reach it only as a JSON field. With the flag off, every tool definition and every
answer is byte for byte what it was before structured output existed. See [Configuration](/en/docs/configuration/).

Independently of this setting, every project's indexed documents are also MCP **resources**
(`contextator://<project>/<source>/<path>`), listed and read behind the same auth and never reaching
beyond what `read_document` itself can reach.

## What the tools cannot do

- Write, edit or delete anything.
- See another project's documents.
- Read files that were not indexed.
- Trigger an index run. Indexing is an operator action — dashboard, API or webhook.
