# Content Types

> The four content-type flavors a source can pick — plain, Obsidian vault, Notion export and OpenAPI/Swagger — and what each one transforms before chunking.

- Updated: 2026-09-25
- Source: https://contextator.com/en/docs/content-types/
- Language: en-US
- Author: Muhammet Şafak

---
Documentation exported from another tool is Markdown, but carries that tool's own syntax. A source's
**content type** — sometimes called its flavor — applies a small transform so an agent reads plain
Markdown regardless of where a file came from. Your files are never modified: not the mounted folder,
not the git checkout, not the uploaded tree.

## Every file becomes Markdown first — one exception

Before a content type runs at all, every file a source reads is converted to Markdown by its own file
type: `.html`/`.htm`, `.docx`, `.pdf` and `.csv` are each converted; `.md`, `.mdx` and `.txt` pass
through unchanged, already being text. This is what lets a Confluence page or a crawled
documentation-site page arrive as Markdown too — see [Confluence](/en/docs/confluence/) and
[Documentation Site Source](/en/docs/documentation-site-source/). The content type below is a *second*
transform, applied to that Markdown, and depends on where the files came from rather than what format
they were in.

The one exception is **OpenAPI / Swagger**: a `.yaml`, `.yml` or `.json` specification under that
content type is never converted to Markdown at all. It is read as structured data and rendered straight
into one document per operation — see below.

## The four content types

| Content type | Use it for | What it does |
|---------------|-----------|---------------|
| **Plain Markdown / text** | Everything normal — the default | Nothing |
| **Obsidian vault** | An Obsidian vault, however it arrived | Rewrites `[[wikilinks]]`, callouts and comments |
| **Notion export** | A Notion *Export → Markdown & CSV* zip | Strips the 32-character page id Notion appends to file and folder names, and fixes the links pointing at them |
| **OpenAPI / Swagger** | A source holding API specifications | Reads `.yaml`, `.yml` and `.json` as specifications and turns one file into one document per operation |

**Add source** picks this per source: the dashboard's **Obsidian vault** tab selects the Obsidian
content type for you; a Notion export zip needs **Upload files** with **Notion export** chosen
explicitly, since the tab does not imply it; **OpenAPI / Swagger** must always be chosen deliberately.
See [Document Sources](/en/docs/document-sources/).

## Obsidian vault

Every wikilink form becomes a standard Markdown link, `> [!NOTE]` becomes `> **Note:**` so the callout
kind stays searchable in plain text, and `%%comments%%` are removed. The full table of rewrites is on
[Obsidian Vaults](/en/docs/obsidian-vaults/).

## Notion export

A Notion export zip names files and folders like this:

```
Getting started 1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d.md
Engineering 9f8e7d6c5b4a39281706f5e4d3c2b1a0/Runbook 0a1b2c3d….md
```

With **Notion export** selected, the 32-character id is removed from every path segment and from the
links that point at those files, so documents read as:

```
Getting started.md
Engineering/Runbook.md
```

For the live API instead of an export zip, use a [Notion](/en/docs/notion/) source instead — it needs
no content type, since there is no id to strip from an API response.

## OpenAPI / Swagger

This is the one content type that does something a transform alone cannot: it says a file is **not one
document**. A specification is indexed as **one document per `(method, path)` operation** — path,
method, summary, parameters, request and response schemas and examples, each under its own heading, so
a search hit carries a breadcrumb like `GET /pets/{petId} > Responses > 200` and `search_docs` returns
the endpoint, not the file.

- **`.yaml`, `.yml` and `.json` are selectable only under this content type.** No extension implies a
  specification on its own — a `.yaml` might just as well be a Helm values file or a CI config — so
  choosing this content type is what makes those extensions readable at all, on that source alone.
- **Each document is stored at `<file>/<method>-<path>`**, e.g. `api/petstore.yaml/get-pets-petId`. The
  path comes from the method and the URL path alone, so re-indexing the same specification, even
  reformatted, lands on the same documents.
- **`$ref` is resolved within the file.** A recursive schema renders until it points back at itself and
  says so; anything past eight levels deep says it stopped there. A reference into another file is
  named rather than followed.
- **Markdown beside the specifications stays Markdown** — a `README.md` in the same source is still one
  ordinary document.
- **A file that is not a valid specification is refused by name**, the same way an unconvertible PDF
  is — the reason is shown on the source's row, and the rest of the source indexes normally. So is one
  written to make a renderer fail, such as a self-referential example or a `$ref` that will not decode:
  nothing a specification can contain fails the whole run.
- **Swagger 2.0 is read as well as OpenAPI 3**, `definitions`, `in: body` parameters and
  `host`/`basePath` included; a 3.1 path item that is itself a `$ref` is followed.
- **A specification is measured against `MAX_SPEC_FILE_BYTES` (8 MiB) before it is read** — its own
  ceiling, well below the one for converted files, because parsing one produces an object graph around
  fifty-five times the size of the file, held in memory for as long as that file is being indexed. At
  the default ceiling that is roughly **400 MB of heap** while one specification indexes; lower the
  variable on a container that cannot spare it. See [Configuration](/en/docs/configuration/).
- **One rendered document is capped at 2 000 lines, and one specification at 5 000 operations.** Neither
  is reachable by a real API — the largest published specifications run to roughly a thousand
  operations — and both exist because the file-size ceiling bounds the parse, not what a file asks to
  be rendered into.
- Two releases of the same API in one project do not collide: give each specification's source a
  [version](/en/docs/document-sources/), and `v2/openapi.yaml/get-pets` and `v3/openapi.yaml/get-pets`
  are separate documents that `search_docs`'s `version` filter tells apart.

## Changing the content type later

Editing a source and picking another content type has one subtlety worth knowing: Contextator decides
what to re-embed by hashing each file's **raw bytes**, before either transform runs. Changing the
content type alone would never touch a file whose bytes did not change — so instead, changing it
**drops that source's stored hashes and queues an index run**, and every file is re-processed with the
new transform. The practical consequence is that changing the content type of a large source costs a
full re-embed of it; changing anything else does not.

## Picking the right one

- If in doubt, leave a source on **Plain Markdown / text** — the other transforms only help when their
  syntax is actually present, and never improve ordinary Markdown.
- A Notion export zip needs **Notion export** chosen explicitly on an **Upload** source; nothing infers
  it from the file names.
- **OpenAPI / Swagger** must be chosen deliberately, and it is what makes `.yaml`, `.yml` and `.json`
  selectable in **File types** at all on that source.
