Contextator
ENTR

Operating

MCP Tools

The three read-only tools every project publishes — search_docs, list_topics and read_document — their arguments, their limits, and what they cannot do.

Updated:

Every project publishes the same three tools, scoped to its own documents. They are read-only: an agent can search and read, never write.

When a client connects, the server also sends instructions naming the project, how many documents and chunks it holds, and when to use which tool — which is why you can simply ask a question instead of telling the agent which tool to call.


search_docs

Hybrid search over the project’s chunks: meaning and exact wording are searched at once and fused by rank, so HALYARD_DISPATCH_TIMEOUT finds its page as readily as a question does.

Argument Type Default
query string (1–2000 characters) required
limit integer 1–20 5
source string — one source name, the first segment of the paths list_topics shows every source
path_prefix string — e.g. handbook/operations the whole project
version string — an exact release label as the project set it, e.g. v3 every version

The last three narrow the search and are all optional; omitting them searches everything, as it always did. An unknown source or version is answered with the ones this project actually has, never with an empty page — there is no ordering and no latest.

Returns ranked excerpts, each with its file path, heading breadcrumb and similarity score:

Found 3 results for "how do I rotate the signing key" in project "handbook":

### 1. runbooks/keys.md — Runbooks > Signing keys > Rotation (score 0.871)
Rotate the signing key by generating a new pair with `keytool`, publishing the public
half to the JWKS endpoint, and keeping the previous key active for 24 hours…

### 2. security/policy.md — Security > Key material (score 0.842)
…

Because the breadcrumb is part of what gets embedded, a question phrased like a heading tends to find the right section directly.

Each excerpt is rendered with the chunk before and after it, marked with a leading and trailing …, so an agent usually does not have to spend a read_document call to see the sentence a chunk boundary cut in half. At most two excerpts come from any one document, because five results that are five consecutive chunks of one page answer the question once and crowd out four other pages. The whole answer is capped at SEARCH_MAX_RESULT_CHARS and says so when it cuts — see Configuration.

Special answers instead of results:

  • The project has nothing indexed yet → a message saying so, and to index it.
  • The project was indexed with a different embedding model → an error asking you to re-index, rather than confidently wrong matches.
  • Nothing clears the relevance floor, a cosine similarity of 0.82 by default → no good match and a pointer at list_topics, rather than the least bad hit it found. See Embedding Models for what the numbers mean.

list_topics

Argument Type Default
cursor string — the next_cursor a previous call ended with start of the listing
limit integer 1–1000 200

One call returns limit documents (200 unless said otherwise, at most 1000). Returns the page grouped by directory:

Project "handbook": 84 documents, 412 chunks
Sources (the first path segment): policies (local), api (git: API reference), notion (notion)

policies/ — 12 documents, 61 chunks
  • policies/onboarding.md — Onboarding (9 chunks)
  • policies/travel.md — Travel policy (4 chunks)

api/reference/ — 31 documents, 180 chunks
  • api/reference/auth.md — Authentication (12 chunks)
  …

Useful to an agent that wants to know what exists before searching, and to you when you want to check what was actually indexed. When more documents remain past the page, the answer ends with a next_cursor value and a line telling the agent to call list_topics again with that value as cursor.

read_document

Argument Type Default
path string — exactly as shown by search_docs or list_topics required
heading string — a heading breadcrumb out of a search result, e.g. Guide > Install > Docker the whole document
from, to integer — a chunk range, 0-based and inclusive the whole document
max_tokens integer 200–20000 4000

Returns the Markdown of one indexed file:

File: api/reference/auth.md
Title: Authentication

---

# Authentication
…

Pass heading with a breadcrumb from a search result to read only that section and the subsections under it, or from/to for a chunk range — either is far cheaper than a whole page.

Only paths that were indexed for that project are served — never an arbitrary filesystem path. Output is capped at max_tokens, counted with the embedding model’s own tokenizer, and says where it cut and how to ask for the rest.

Reading a document is not the same as reading a file. read_document serves the text this server indexed, stored beside the chunks, and never touches the filesystem. So a document stays readable after its file is renamed, its git checkout is re-cloned, or the whole source directory is unmounted, and a page that was never indexed can never be served by mistake. The text is stored as the indexer transformed it, which is what keeps read_document and search_docs from ever disagreeing about what a page says; the price is that an Obsidian note’s original [[wikilink]] syntax is not readable through MCP, only the Markdown link it became. Documents indexed by an older version hold no stored text and are read from disk until their next index run, which is the one case where a missing file is still an error.


Paths and sources

Every path begins with the name of the source it came from:

handbook/install.md      ← source "handbook"
api-repo/install.md      ← source "api-repo"

That is what tells you where an answer came from, and what lets two sources hold the same file name. See Document Sources.

Getting better answers

  • Ask questions, but identifiers work too. Search is hybrid, so an environment variable, a header or an error code finds its page by exact wording; for everything else a full question carries more signal for the meaning half than two words do.
  • Ask in the language the page is written in. The default model covers 100 languages including Turkish, but retrieval does not cross between them — that is a measured limit of the model, not a switch.
  • Narrow it when you know where the answer lives. source and path_prefix cut the search to one source or one directory, and version to one release, so v2 and v3 of the same documentation do not answer for each other.
  • Raise limit when a topic is spread over many files — an agent can ask for up to 20 excerpts.
  • Structure your documentation with headings. Chunks are cut at headings and carry their breadcrumb, so good headings directly improve retrieval.
  • Keep unrelated bodies of knowledge in separate projects. Precision drops when one endpoint mixes a handbook, an API reference and a journal.
  • Tell the agent to cite the file path. The server already asks it to; it is worth reinforcing in your own prompt.

Structured output

Set MCP_STRUCTURED_OUTPUT=1 and every tool also publishes an outputSchema and returns structuredContent alongside the same text (MCP 2025-06-18). It is off by default: a client that reads structured content can stop reading the text, and Claude Code does exactly that once a result carries structuredContent (anthropics/claude-code#55677, #79944), so the prose guiding the model on how to read excerpts would reach it only as a JSON field. With the flag off, every tool definition and every answer is byte for byte what it was before structured output existed. See Configuration.

Independently of this setting, every project’s indexed documents are also MCP resources (contextator://<project>/<source>/<path>), listed and read behind the same auth and never reaching beyond what read_document itself can reach.

What the tools cannot do

  • Write, edit or delete anything.
  • See another project’s documents.
  • Read files that were not indexed.
  • Trigger an index run. Indexing is an operator action — dashboard, API or webhook.

Arrow keys to move, Enter to open.