by pvliesdonk
Offers full‑text (FTS5) and semantic vector search over a directory of Markdown files, with front‑matter‑aware indexing, incremental reindexing, and read/write tools exposed through the Model Context Protocol.
Markdown Vault MCP turns any folder of Markdown notes (Obsidian vaults, Zettelkasten, PARA structures, etc.) into a searchable knowledge base. It indexes the content, YAML front‑matter, and attachments, then serves search, read, write, and edit operations via the Model Context Protocol for LLMs and other clients.
pip install markdown-vault-mcp
# or Docker
docker pull ghcr.io/pvliesdonk/markdown-vault-mcp:latest
MARKDOWN_VAULT_MCP_SOURCE_DIR to point at your vault directory. Optional variables control embedding provider, git integration, authentication, etc.export MARKDOWN_VAULT_MCP_SOURCE_DIR=/path/to/vault
markdown-vault-mcp serve
For HTTP deployment behind a reverse proxy, use --transport http and configure MARKDOWN_VAULT_MCP_HTTP_PATH or MARKDOWN_VAULT_MCP_BASE_URL.search, read, write, edit, move_folder, or use the built‑in SPA via browse_vault.Q: Do I need a separate scheduler for indexing? A: No. The server performs incremental reindexing on startup and watches the filesystem (or reacts to Git/webhook events) to keep the index up‑to‑date.
Q: Which embedding providers are supported?
A: FastEmbed (local), Ollama, and OpenAI. The provider is selected via MARKDOWN_VAULT_MCP_EMBEDDING_PROVIDER or auto‑detected from available credentials.
Q: Can I run the server without write permissions?
A: Yes. Set MARKDOWN_VAULT_MCP_READ_ONLY=true (default) to expose only read‑only tools.
Q: How does authentication work?
A: For HTTP transports you can enable a static bearer token (MARKDOWN_VAULT_MCP_BEARER_TOKEN) or full OIDC (set the four OIDC variables). Both can be active simultaneously.
Q: How are large documents handled in search results?
A: Results are capped to a configurable number of chunks per document (MARKDOWN_VAULT_MCP_CHUNKS_PER_FILE). Longer documents are down‑weighted and split adaptively at heading levels.
Q: Is there a UI to explore the vault?
A: Yes. The built‑in SPA (ui://markdown_vault_mcp/app.html) provides a vault browser, graph explorer, context card, and note preview, accessible via MCP Apps clients.
A generic markdown vault MCP server with FTS5 full-text search, semantic vector search, frontmatter-aware indexing, incremental reindexing, and non-markdown attachment support.
Documentation | Config wizard | PyPI | Docker
Point it at a directory of Markdown files (an Obsidian vault, a docs folder, a Zettelkasten, a PARA vault) and it exposes search, read, write, and edit tools over the Model Context Protocol.
read(path, section=heading)Upgrading. As of this release,
searchreturns query-relevant snippets in thecontentfield by default (approximately 200 words). Passsnippet_words=0to recover the prior full-chunk behaviour, or useread(path, section=heading)to fetch the full section after seeing a snippet. Documents are also re-chunked on nextreindexto honour the adaptiveMARKDOWN_VAULT_MCP_MAX_CHUNK_WORDSthreshold (default 400).
okf_version declaration in the root index.md) and annotates search/read results with each note's type, lifecycle status, staleness, and trust tier. Those dimensions are filterable and also nudge ranking on a detected bundle (deprecated and stale notes rank lower, and the reserved index.md / log.md navigation files are demoted below real notes), and the server adds an okf_validate conformance audit plus one-shot migration transforms (okf_convert_links, okf_generate_index, okf_seed_log) for moving a vault into the format, plus a downloadable bundle export served through create_download_link with an okf-bundle reference. Read semantics are controlled by MARKDOWN_VAULT_MCP_OKF_MODE_conventions.md files carry your authoring rules, such as "reference notes stay self-contained"; the server surfaces them to LLM clients at write time via the get_conventions tool and in write/edit results, without interpreting themGIT_ASKPASSmove_folder, git history, manual git sync, one-time transfer links, and admin operations; plus 6 app-only tools for MCP Apps clientsWith this server mounted in Claude, you can:
3-Resources/, and link any existing notes on the topic." Claude composes fetch + search + write.write with wikilinks. See the Research workflows guide for the full loop.conversation_search + recent_chats + write. The para-capture-chats prompt is the one-click version.propose-links prompt from the + menu: it scans recently modified notes and proposes links between notes that aren't yet connected, writing them on confirmation.<existing note> instead of duplicating." Claude composes read + write + delete.The vault needs no external scheduler or separate capture app: it sits behind your conversations and absorbs their output.
pip install markdown-vault-mcp
With optional dependencies:
pip install markdown-vault-mcp[mcp] # FastMCP server
pip install markdown-vault-mcp[embeddings-api] # Ollama/OpenAI embeddings via API
pip install markdown-vault-mcp[embeddings] # FastEmbed local embeddings
pip install markdown-vault-mcp[all] # MCP + FastEmbed + API embeddings
git clone https://github.com/pvliesdonk/markdown-vault-mcp.git
cd markdown-vault-mcp
uv sync --all-extras --all-groups
docker pull ghcr.io/pvliesdonk/markdown-vault-mcp:latest
The Docker image uses [all] (MCP + FastEmbed + API embeddings). By default, semantic search works locally with FastEmbed and can switch to Ollama/OpenAI when configured. A compose.yml ships at the repo root as a starting point: copy .env.example to .env, edit, and docker compose up -d.
To attach a remote Python debugger (development only; the protocol is unauthenticated), see Remote debugging.
Download .deb or .rpm packages from the GitHub Releases page. Both install a hardened systemd unit; env configuration is sourced from /etc/markdown-vault-mcp/env (copy from the shipped /etc/markdown-vault-mcp/env.example). See the systemd deployment guide for details.
Download the .mcpb bundle from the GitHub Releases page. Double-click to install, or run:
mcpb install markdown-vault-mcp-<version>.mcpb
Claude Desktop opens a GUI wizard that prompts for required env vars; no manual JSON editing is needed. See Step 0 of the Claude Desktop guide for details.
/plugin marketplace add pvliesdonk/claude-plugins
/plugin install markdown-vault-mcp@pvliesdonk
Installs the MCP server and the vault-workflow skill. See the Claude Code plugin guide for details.
from pathlib import Path
from markdown_vault_mcp.vault import Vault
vault = Vault(source_dir=Path("/path/to/vault"))
vault.index.build_index()
results = vault.reader.search("query text", limit=10)
export MARKDOWN_VAULT_MCP_SOURCE_DIR=/path/to/vault
markdown-vault-mcp serve
Copy an example env file:
cp examples/obsidian-readonly.env .env
Edit .env to set MARKDOWN_VAULT_MCP_SOURCE_DIR to the absolute path of your vault on the host.
Start the service:
docker compose up -d
Check the logs:
docker compose logs -f markdown-vault-mcp
| File | Description |
|---|---|
examples/obsidian-readonly.env |
Obsidian vault, read-only, Ollama embeddings |
examples/obsidian-readwrite.env |
Obsidian vault, read-write with git auto-commit |
examples/obsidian-oidc.env |
Obsidian vault, read-only, OIDC authentication (Authelia) |
examples/ifcraftcorpus.env |
Strict frontmatter enforcement, read-only corpus |
For reverse proxy (Traefik) and deployment setup, see docs/deployment.md.
The server registers a built-in get_server_info tool (via fastmcp_pvl_core.register_server_info_tool) so operators can confirm the deployed version with a single MCP call. The response carries server_name, server_version, and core_version.
All configuration is via environment variables with the MARKDOWN_VAULT_MCP_ prefix (except embedding provider settings, which use their own conventions).
Domain environment variables use the MARKDOWN_VAULT_MCP_ prefix (except the
embedding-provider conventions OLLAMA_HOST, OPENAI_API_KEY,
OPENAI_BASE_URL, and OPENAI_EMBEDDING_MODEL, which are also honoured bare):
| Variable | Default | Required | Description |
|---|---|---|---|
OLLAMA_HOST |
http://localhost:11434 |
No | Ollama server URL for the ollama embedding provider. Bare (not MARKDOWN_VAULT_MCP_-prefixed), matching the Ollama ecosystem convention. |
OPENAI_API_KEY |
(none) | No | OpenAI API key for the openai embedding provider, and the fallback key for the summarize tool when MARKDOWN_VAULT_MCP_SUMMARIZE_OPENAI_API_KEY is unset. Bare (not MARKDOWN_VAULT_MCP_-prefixed), matching the OpenAI ecosystem convention. |
OPENAI_BASE_URL |
(none) | No | Bare fallback for MARKDOWN_VAULT_MCP_OPENAI_BASE_URL (embeddings). For the summarize tool it only routes traffic when an API key already enables the feature; it never enables summarize by itself. |
OPENAI_EMBEDDING_MODEL |
(none) | No | Bare fallback for MARKDOWN_VAULT_MCP_OPENAI_EMBEDDING_MODEL. |
MARKDOWN_VAULT_MCP_BUILD_TIMEOUT_S |
60 |
No | Maximum seconds an index-backed tool or resource waits for the FTS index to become queryable during a cold-start background build before raising IndexUnavailableError(reason="timeout"). Increase for large vaults. |
MARKDOWN_VAULT_MCP_DRAIN_TIMEOUT_S |
60 |
No | Maximum seconds an index-querying read tool waits for the IndexWriter to drain when called with wait_for_pending_writes=true. On timeout the tool answers from the current index and reports index_stale=true in the response _meta. |
MARKDOWN_VAULT_MCP_SOURCE_DIR |
/data/vault |
No | Path to the markdown vault directory. Required; the server refuses to start without it. Symbolic links inside the vault are followed on Python 3.13+. |
MARKDOWN_VAULT_MCP_READ_ONLY |
true |
No | Set to false to enable write tools (write, edit, delete, rename). |
MARKDOWN_VAULT_MCP_DISABLE_APPS_UI |
false |
No | Hide the MCP Apps UI tools (browse_vault, show_context) from the tool listing for clients that do not render MCP Apps panels. |
MARKDOWN_VAULT_MCP_INDEX_PATH |
(none) | No | Path to the SQLite FTS5 index file; unset keeps the index in memory. Set it for persistence across restarts. |
MARKDOWN_VAULT_MCP_STATE_PATH |
(none) | No | Path to the change-tracking state file. Defaults to {SOURCE_DIR}/.markdown_vault_mcp/state.json. |
MARKDOWN_VAULT_MCP_EMBEDDINGS_PATH |
(none) | No | Path to the numpy embeddings file; required to enable semantic search. |
MARKDOWN_VAULT_MCP_INDEXED_FIELDS |
(none) | No | Comma-separated frontmatter fields promoted to the tag index for structured filtering. Changing it cold-rebuilds the index once on next startup; SEARCHABLE_FIELDS inherits this value when unset. |
MARKDOWN_VAULT_MCP_REQUIRED_FIELDS |
(none) | No | Comma-separated frontmatter fields required on every document; documents missing any are excluded from the index. |
MARKDOWN_VAULT_MCP_EXCLUDE |
(none) | No | Comma-separated glob patterns excluded from scanning (.obsidian/,.trash/). |
MARKDOWN_VAULT_MCP_TITLE_FIELD |
title |
No | Frontmatter field used as the document title (falls back to title, the first H1, then the filename). Changing it cold-rebuilds the index once on next startup. |
MARKDOWN_VAULT_MCP_SEARCHABLE_FIELDS |
(none) | No | Comma-separated frontmatter fields whose text values become keyword-searchable and enrich first-chunk embeddings. Inherits INDEXED_FIELDS when unset; the sentinel none means filterable but not searchable. Changing it cold-rebuilds the index and re-embeds once on next startup. |
MARKDOWN_VAULT_MCP_TEMPLATES_FOLDER |
_templates |
No | Relative folder where note templates live (used by the create_from_template prompt). |
MARKDOWN_VAULT_MCP_PROMPTS_FOLDER |
(none) | No | Directory of .md prompt files that extend or override built-in prompts; a relative path is resolved against SOURCE_DIR. |
MARKDOWN_VAULT_MCP_CONVENTIONS_FILE |
_conventions.md |
No | Filename of the per-folder conventions files surfaced to clients at write time (bare .md filename without glob characters). Set to none to disable folder conventions. |
MARKDOWN_VAULT_MCP_OKF_MODE |
auto |
No | OKF (Open Knowledge Format) read semantics. With auto (the default), read annotations switch on when the vault declares an OKF version in its root index.md. Use off to disable OKF semantics entirely, or on to force them for an undeclared vault. Annotations are read-only; write behavior is never affected. |
MARKDOWN_VAULT_MCP_OKF_WRITE |
false |
No | OKF (Open Knowledge Format) enforced write layer. When true on an OKF-active vault, the server stamps generated provenance on each write and clears any verified attestation when a note's content changes. It also keeps each written folder's log.md and index.md current, and exposes the okf_verify tool. Requires OKF_MODE to be auto or on (a true value with OKF_MODE=off is a config error). Off by default. |
MARKDOWN_VAULT_MCP_ATTACHMENT_EXTENSIONS |
(none) | No | Comma-separated allowed attachment extensions without the dot (such as pdf,png,jpg); use * to allow every non-markdown file. Unset selects the built-in allowlist. |
MARKDOWN_VAULT_MCP_MAX_ATTACHMENT_SIZE_MB |
1.0 |
No | Maximum attachment size in MB returned by read / accepted by write; 0 disables the limit. |
MARKDOWN_VAULT_MCP_MAX_NOTE_READ_BYTES |
262144 |
No | Maximum bytes returned by a full-document read of a note; use read(path, section=…) for partial reads. 0 disables the limit. |
MARKDOWN_VAULT_MCP_CHUNKS_PER_FILE |
2 |
No | Maximum chunks returned per document in search results. |
MARKDOWN_VAULT_MCP_SNIPPET_WORDS |
200 |
No | Width of the snippet window (words) in search results; 0 returns full chunk content. |
MARKDOWN_VAULT_MCP_LENGTH_DOWNWEIGHT_ALPHA |
0.25 |
No | Down-weights longer chunks in ranking: score / (1 + alpha * log(chunk_count)). |
MARKDOWN_VAULT_MCP_MAX_CHUNK_WORDS |
400 |
No | Word cap per chunk; the adaptive chunker splits at deeper heading levels, then paragraph/word boundaries, to respect it. Match it to the embedding model's context. A reindex applies a new value. |
MARKDOWN_VAULT_MCP_MAX_CHUNK_CHARS |
(none) | No | Character cap enforced alongside MAX_CHUNK_WORDS to bound token-dense chunks. Unset derives min(1500, model context * 2.8). Set a positive value for an exact cap, or -1 to scale with the model's full context (can exhaust memory on long-context models). A reindex applies a new value. |
MARKDOWN_VAULT_MCP_CHUNK_OVERLAP_WORDS |
40 |
No | Words of overlap between adjacent budget-split fragments of the same heading section (0 disables). A reindex applies a new value. |
MARKDOWN_VAULT_MCP_FOLDER_WEIGHTS |
(none) | No | Folder-prefix score multipliers (prefix:weight pairs, comma-separated, weights > 0) applied to all search modes; the deepest matching prefix wins (sessions:0.5 demotes sessions/**). |
MARKDOWN_VAULT_MCP_FTS_WEIGHTS |
(none) | No | Per-column BM25 weights (column:weight pairs, comma-separated, weights >= 0) for keyword ranking. Columns: path, title, folder, heading, content, summary. |
MARKDOWN_VAULT_MCP_EMBEDDING_PROVIDER |
(none) | No | Embedding provider: openai, ollama, or fastembed. Unset auto-detects from the environment. |
MARKDOWN_VAULT_MCP_OLLAMA_MODEL |
nomic-embed-text |
No | Ollama embedding model name. |
MARKDOWN_VAULT_MCP_OLLAMA_CPU_ONLY |
false |
No | Force Ollama to embed on CPU only. |
MARKDOWN_VAULT_MCP_OPENAI_BASE_URL |
https://api.openai.com/v1 |
No | OpenAI-compatible API base URL for embeddings; the bare OPENAI_BASE_URL is honoured as a fallback. |
MARKDOWN_VAULT_MCP_OPENAI_EMBEDDING_MODEL |
text-embedding-3-small |
No | OpenAI-compatible embedding model name; the bare OPENAI_EMBEDDING_MODEL is honoured as a fallback. |
MARKDOWN_VAULT_MCP_FASTEMBED_MODEL |
BAAI/bge-small-en-v1.5 |
No | FastEmbed model name. |
MARKDOWN_VAULT_MCP_FASTEMBED_CACHE_DIR |
(none) | No | FastEmbed model cache directory (in Docker, stored under /data/state/fastembed). |
MARKDOWN_VAULT_MCP_EMBED_CONTEXT |
false |
No | Enrich embedding input with the note title, chunk heading, and (first chunk) searchable-field values. Flipping it re-embeds the whole vault once on next startup. |
MARKDOWN_VAULT_MCP_GIT_TOKEN |
(none) | No | Token/password for HTTPS git auth; remotes must be HTTPS when set. |
MARKDOWN_VAULT_MCP_GIT_REPO_URL |
(none) | No | HTTPS remote URL for managed git mode: the server clones into an empty SOURCE_DIR on startup (or validates an existing origin) and enables the pull loop, auto-commit, and deferred push. |
MARKDOWN_VAULT_MCP_GIT_USERNAME |
x-access-token |
No | Username for HTTPS git auth prompts (x-access-token for GitHub, oauth2 for GitLab, the account name for Bitbucket). |
MARKDOWN_VAULT_MCP_GIT_PULL_INTERVAL_S |
600 |
No | Seconds between git fetch + fast-forward update attempts; 0 disables periodic pull. |
MARKDOWN_VAULT_MCP_GIT_PUSH_DELAY_S |
30.0 |
No | Seconds of write-idle time before pushing; 0 pushes only on shutdown. |
MARKDOWN_VAULT_MCP_GIT_COMMIT_NAME |
markdown-vault-mcp |
No | Git committer name for auto-commits; set this in Docker where git config user.name is empty. |
MARKDOWN_VAULT_MCP_GIT_COMMIT_EMAIL |
noreply@markdown-vault-mcp |
No | Git committer email for auto-commits. |
MARKDOWN_VAULT_MCP_GIT_COMMIT_NAME_CLAIM |
(none) | No | OIDC claim key used as the commit author name (such as name); overrides GIT_COMMIT_NAME per request when an OIDC token is present. |
MARKDOWN_VAULT_MCP_GIT_COMMIT_EMAIL_CLAIM |
(none) | No | OIDC claim key used as the commit author email (such as email); overrides GIT_COMMIT_EMAIL per request when an OIDC token is present. |
MARKDOWN_VAULT_MCP_GIT_LFS |
true |
No | Run git lfs pull on startup to fetch LFS-tracked attachments; set to false for repos without LFS. |
MARKDOWN_VAULT_MCP_FILE_WATCHER |
true |
No | Watch the vault for external filesystem changes; auto-disabled when git pull or the webhook is active. Requires the file-watcher extra. |
MARKDOWN_VAULT_MCP_FILE_WATCHER_DEBOUNCE_S |
2.0 |
No | Seconds of quiet after the last filesystem event before reindexing. |
MARKDOWN_VAULT_MCP_FILE_WATCHER_ROOT_FLOOR |
true |
No | Keep the non-recursive watch on the vault root; set false to register zero source-dir-rooted FSEvents streams (avoids repeated macOS access prompts on a home-rooted vault) at the cost of root-level files relying on scans. |
MARKDOWN_VAULT_MCP_GITHUB_WEBHOOK_SECRET |
(none) | No | Shared secret for the GitHub push-event webhook; when set, mounts POST /github-webhook on HTTP/SSE transports to trigger an immediate pull + reindex on push events. |
MARKDOWN_VAULT_MCP_SUMMARIZE_PROVIDER |
(none) | No | Summarization backend (only openai is recognised). Unset auto-detects: the backend activates when credentials or an explicit endpoint are present. |
MARKDOWN_VAULT_MCP_SUMMARIZE_OPENAI_API_KEY |
(none) | No | API key for the OpenAI-compatible summarize endpoint; the bare OPENAI_API_KEY is honoured as a fallback. Unset works for keyless local endpoints (Ollama). |
MARKDOWN_VAULT_MCP_SUMMARIZE_OPENAI_BASE_URL |
(none) | No | OpenAI-compatible endpoint base URL for the summarize tool; setting it enables the tool even without an API key. The bare OPENAI_BASE_URL routes traffic only when a key already enables the feature. |
MARKDOWN_VAULT_MCP_SUMMARIZE_OPENAI_MODEL |
gpt-5-mini |
No | Chat model id used for summaries. |
MARKDOWN_VAULT_MCP_SUMMARIZE_MAX_TOKENS |
8192 |
No | Upper bound on generated tokens per summarize call; on reasoning models this budget also covers internal reasoning tokens. |
MARKDOWN_VAULT_MCP_SUMMARIZE_MAX_NOTES |
50 |
No | Cap on the number of notes summarised in one call (subtree expansion truncates to this many). |
MARKDOWN_VAULT_MCP_SUMMARIZE_MAX_INPUT_CHARS |
200000 |
No | Aggregate cap on note characters sent to the model in one call; excess is truncated with a flag on the result. |
MARKDOWN_VAULT_MCP_SUMMARIZE_TIMEOUT |
120.0 |
No | Per-request wall-clock budget in seconds for a single summarize backend call; keep it below the MCP client's request timeout so the server-side error wins the race. |
MARKDOWN_VAULT_MCP_SUMMARIZE_INLINE_TIMEOUT |
30.0 |
No | Soft deadline in seconds before a still-running summarize call is promoted to a background job retrievable via get_summary. Must be <= SUMMARIZE_TIMEOUT. |
MARKDOWN_VAULT_MCP_TRANSFER_TTL_DEFAULT_S |
3600.0 |
No | Link lifetime in seconds when the caller requests no explicit TTL. |
MARKDOWN_VAULT_MCP_TRANSFER_TTL_MAX_S |
86400.0 |
No | Ceiling in seconds a caller-requested link TTL is clamped to. |
MARKDOWN_VAULT_MCP_TRANSFER_GRACE_TTL_S |
60.0 |
No | Post-success grace window in seconds: a served token's TTL shrinks to this so a stalled transfer can retry within it. |
MARKDOWN_VAULT_MCP_TRANSFER_LEASE_S |
60.0 |
No | Crashed-handler reclaim window in seconds for an in-flight reservation. |
MARKDOWN_VAULT_MCP_TRANSFER_MAX_UPLOAD_BYTES |
104857600 |
No | Maximum size in bytes of a single upload. |
Domain-config fields are composed inside src/markdown_vault_mcp/config.py between the CONFIG-FIELDS-START / CONFIG-FIELDS-END sentinels; env reads go through fastmcp_pvl_core.env(_ENV_PREFIX, "SUFFIX", default) so naming stays consistent, and field invariants go in __post_init__ between the CONFIG-VALIDATE-START / CONFIG-VALIDATE-END sentinels. Each field's metadata help and tags generate the table above directly, so keep them accurate and complete. See the configuration reference for the detailed prose documentation of every variable.
| Variable | Default | Description |
|---|---|---|
MARKDOWN_VAULT_MCP_KV_STORE_URL |
file:///data/state |
Persistent-state backend URL shared by every pvl-core subsystem that needs state. memory:// is in-process and lost on restart; file:///path persists on one server; redis://, dynamodb:// and mongodb:// each need their matching extra. Defaults to file:///data/state when unset. |
FASTMCP_LOG_LEVEL |
INFO |
Log level for FastMCP internals and app loggers (DEBUG / INFO / WARNING / ERROR / CRITICAL). The -v CLI flag overrides to DEBUG. |
FASTMCP_ENABLE_RICH_LOGGING |
true |
Set false for plain or structured JSON log output. |
The chunker's character cap (
MARKDOWN_VAULT_MCP_MAX_CHUNK_CHARS) is derived from the embedding model's context length, so changing the embedding model re-chunks the FTS index (not just the embeddings) and triggers an automatic cold rebuild of the index on the next startup. The defaults stay memory-light (BAAI/bge-small-en-v1.5for FastEmbed,nomic-embed-textfor Ollama); long-context models, such asnomic-ai/nomic-embed-text-v1.5(8192 tokens) for FastEmbed orbge-m3:latestfor Ollama, are opt-in and need substantially more RAM/VRAM during indexing.
Git integration has three modes:
MARKDOWN_VAULT_MCP_GIT_REPO_URL set): server owns repo setup.
On startup it clones into SOURCE_DIR when empty, or validates existing origin.
Pull loop + auto-commit + deferred push are enabled.GIT_REPO_URL): writes are committed to a local git repo if SOURCE_DIR is already a git checkout. The server neither pulls nor pushes.SOURCE_DIR is not a git repo, git callbacks are no-ops.When token auth is used (MARKDOWN_VAULT_MCP_GIT_TOKEN), remotes must be HTTPS.
SSH remotes (such as git@github.com:owner/repo.git) are rejected with a startup error.
Fix with: git -C /path/to/vault remote set-url origin https://github.com/owner/repo.git
Backward compatibility: MARKDOWN_VAULT_MCP_GIT_TOKEN without GIT_REPO_URL still works (legacy mode) but logs a deprecation warning.
Requires the watchdog optional extra: pip install 'markdown-vault-mcp[file-watcher]'. Automatically disabled when GIT_PULL_INTERVAL_S > 0 or GITHUB_WEBHOOK_SECRET is set. The watcher scopes one recursive watch per non-excluded top-level directory (not a single recursive watch on the root), so excluded directories are never registered and content under a deliberately watched dot-directory delivers its own edits. See docs/configuration.md for details.
Non-markdown file support. See Attachments for details.
Simple static token auth for HTTP deployments. Set a single env var; clients must send Authorization: Bearer <token>.
| Variable | Required | Description |
|---|---|---|
MARKDOWN_VAULT_MCP_BEARER_TOKEN |
Yes | Static bearer token; any non-empty string enables auth |
Full OAuth 2.1 authentication for HTTP deployments. OIDC activates when all four required variables are set. See Authentication for setup details.
Multi-auth: If both
BEARER_TOKENand all OIDC variables are set, the server accepts either credential: a valid bearer token or a valid OIDC session. This is useful when different clients use different auth flows (such as Claude web via OIDC and Claude Code via bearer token).
| Variable | Required | Description |
|---|---|---|
MARKDOWN_VAULT_MCP_BASE_URL |
Yes | Public base URL of the server (such as https://mcp.example.com; include prefix if mounted under subpath, such as https://mcp.example.com/vault). Used for OIDC auth and to auto-compute the MCP Apps domain. |
MARKDOWN_VAULT_MCP_OIDC_CONFIG_URL |
Yes | OIDC discovery endpoint (such as https://auth.example.com/.well-known/openid-configuration) |
MARKDOWN_VAULT_MCP_OIDC_CLIENT_ID |
Yes | OIDC client ID registered with your provider |
MARKDOWN_VAULT_MCP_OIDC_CLIENT_SECRET |
Yes | OIDC client secret |
MARKDOWN_VAULT_MCP_OIDC_JWT_SIGNING_KEY |
No | JWT signing key. When unset, derived deterministically from the OIDC client secret, so tokens survive restarts; rotating the secret invalidates issued tokens. Set explicitly (generate with openssl rand -hex 32) to decouple token validity from secret rotation |
MARKDOWN_VAULT_MCP_OIDC_AUDIENCE |
No | Expected JWT audience claim; leave unset if your provider does not set one |
MARKDOWN_VAULT_MCP_OIDC_REQUIRED_SCOPES |
No | Comma-separated required scopes; default openid |
MARKDOWN_VAULT_MCP_OIDC_VERIFY_ACCESS_TOKEN |
No | Set true to verify the upstream access token as a JWT instead of the id token. Only needed when your provider issues JWT access tokens and you require audience-claim validation on that token. Default: verify the id token (works with all providers, including opaque-token issuers like Authelia) |
markdown-vault-mcp <command> [options]
serveStart the MCP server.
markdown-vault-mcp serve [--transport {stdio|sse|http}] [--host HOST] [--port PORT] [--http-path PATH]
| Flag | Default | Description |
|---|---|---|
--transport |
stdio |
MCP transport: stdio (stdin/stdout, default), sse (Server-Sent Events), http (streamable-HTTP). Use http for Docker with a reverse proxy or when OIDC is enabled. |
--host |
127.0.0.1 |
Bind host for the http transport (ignored for stdio and sse); pass 0.0.0.0 to bind all interfaces inside Docker |
--port |
8000 |
Port for the http transport (ignored for stdio and sse) |
--http-path (alias --path) |
env MARKDOWN_VAULT_MCP_HTTP_PATH or /mcp |
MCP HTTP path for http transport; useful for reverse-proxy subpath mounting (such as /vault/mcp). The legacy --path spelling is still accepted. |
By default, HTTP transport serves MCP on /mcp. You can run it under a subpath:
markdown-vault-mcp serve --transport http --http-path /vault/mcp
Equivalent env-based config:
MARKDOWN_VAULT_MCP_HTTP_PATH=/vault/mcp
For reverse proxies, you can either:
/mcp and use proxy rewrite/strip-prefix middleware./vault/mcp) and route without rewrite.When OIDC is enabled under a subpath, the configuration is different: the subpath goes in BASE_URL only, and HTTP_PATH stays at /mcp. See OIDC subpath deployments.
Then your redirect URI is:
https://mcp.example.com/vault/auth/callback
indexBuild the full-text search index.
markdown-vault-mcp index [--source-dir PATH] [--index-path PATH] [--force]
searchSearch the vault from the CLI.
markdown-vault-mcp search <query> [-n LIMIT] [-m {keyword|semantic|hybrid}] [--folder PATH] [--json]
reindexIncrementally reindex the vault (only processes changed files). When semantic search is configured, the vector index is converged to the updated chunk set: exactly the changed documents are re-embedded and orphaned vectors dropped, never the whole corpus.
markdown-vault-mcp reindex [--source-dir PATH] [--index-path PATH]
| Tool | Description |
|---|---|
search |
Hybrid full-text + semantic search with optional frontmatter filters |
read |
Read a document or attachment by relative path |
write |
Create or overwrite a document or attachment |
edit |
Replace text in a document: exact match, line-range, or scoped match with normalized fallback |
delete |
Delete a document or attachment and its index entries |
rename |
Rename/move a document or attachment, updating all index entries; pass update_links=true to also rewrite backlinks in other notes |
move_folder |
Move an entire folder subtree to a new prefix, rewriting all vault links that point into the moved subtree in one call |
list_documents |
List indexed documents; pass include_attachments=true to also list non-markdown files |
list_folders |
List all folder paths in the vault |
list_tags |
List all unique frontmatter tag values |
reindex |
Force a full reindex of the vault |
stats |
Get vault statistics (document count, chunk count, link health metrics, etc.) |
build_embeddings |
Build or rebuild vector embeddings for semantic search |
embeddings_status |
Check embedding provider and index status |
get_index_status |
Check background FTS build state (queryable / building / failed) |
get_backlinks |
Find all documents that link to a given document |
get_outlinks |
Find all links from a document, with existence check |
get_broken_links |
Find all links pointing to non-existent documents |
get_similar |
Find semantically similar notes by document path |
get_toc |
Heading outline for a note or a folder subtree |
get_recent |
Get the most recently modified notes |
get_context |
Get a consolidated context dossier for a note (backlinks, outlinks, similar, folder peers, tags, modified time) |
get_orphan_notes |
Find all notes with no inbound or outbound links |
get_most_linked |
Find the most-linked-to notes ranked by backlink count |
get_connection_path |
Find the shortest path between two notes via BFS on the undirected link graph (max 10 hops) |
summarize |
Summarize a note, a set of notes, or a folder subtree with an LLM; the synthesis references the individual source notes by path. A slow summary is promoted to a background job (retrieved via get_summary) so the tool never hangs. Hidden unless an OpenAI-compatible backend is configured (OPENAI_API_KEY or a base URL). Sends note content to the model provider. |
get_summary |
Retrieve a summary that summarize promoted to a background job, by its job_id. Registered alongside summarize. |
get_history |
List commits that touched a note, attachment, or the whole vault (git-backed vaults only) |
get_diff |
Return a diff of a note or attachment between a reference commit/timestamp and HEAD; binary attachments return a --stat size summary instead of a unified patch (git-backed vaults only) |
git_sync |
Force an immediate git pull / push / both, bypassing the periodic loops. Returns structured state (SHAs, commit counts, Syncthing-style conflict file paths if any). Hidden when MARKDOWN_VAULT_MCP_GIT_REPO_URL isn't set or READ_ONLY=true. |
fetch |
Download a file from a URL and save it to the vault as a note or attachment (MCP-to-MCP transfer) |
create_download_link |
Mint a one-time capability URL to download a vault note or attachment (HTTP/SSE only; BASE_URL required) |
create_upload_link |
Mint a one-time capability URL to upload bytes to a fixed vault path (HTTP/SSE only; BASE_URL required; hidden when READ_ONLY=true) |
browse_vault |
Open the vault explorer SPA in a supporting MCP Apps client |
show_context |
Open the Context Card for a specific note in a supporting MCP Apps client |
Write tools (write, edit, delete, rename, move_folder, fetch, git_sync, create_upload_link) are only available when MARKDOWN_VAULT_MCP_READ_ONLY=false. git_sync also requires managed git mode (MARKDOWN_VAULT_MCP_GIT_REPO_URL set).
summarize is registered only when a summarization backend is configured: an OPENAI_API_KEY, or an OpenAI-compatible base URL for local endpoints that need no key. It needs the openai SDK (pip install 'markdown-vault-mcp[summarize]') and sends note content to the model provider. Any OpenAI-compatible endpoint works: OpenAI itself, a local Ollama (http://localhost:11434/v1, no key needed), the Anthropic compatibility endpoint (https://api.anthropic.com/v1), vLLM, and others. See Configuration for provider recipes.
browse_vault and show_context are LLM-visible in all clients; when called in an MCP Apps-capable client they open the interactive SPA. Six additional internal tools (vault_context, vault_list, vault_read, vault_search, vault_graph_neighborhood, vault_graph_hubs) use visibility="app" and are used by the SPA only; they are never visible to the LLM.
MCP resources expose vault metadata as structured JSON that clients can read directly without invoking tools.
| URI | Description |
|---|---|
config://vault |
Current vault configuration (source dir, indexed fields, read-only state, etc.) |
stats://vault |
Vault statistics (document count, chunk count, embedding count, etc.) |
tags://vault |
All frontmatter tag values grouped by indexed field |
tags://vault/{field} |
Tag values for a specific indexed frontmatter field (template) |
folders://vault |
All folder paths in the vault |
toc://vault/{path} |
Table of contents (heading outline) for a specific document (template) |
similar://vault/{path} |
Top 10 semantically similar notes for a document (template) |
recent://vault |
20 most recently modified notes with ISO timestamps |
ui://markdown_vault_mcp/app.html |
Interactive vault explorer SPA for MCP Apps clients |
Prompt templates guide the LLM through multi-step workflows using the vault tools.
| Prompt | Parameters | Description |
|---|---|---|
summarize |
path |
Read a document and produce a structured summary with key themes and takeaways |
research |
topic |
Search for a topic, synthesize findings, and create a new note at research/{topic}.md |
discuss |
path |
Analyze a document and suggest improvements using edit (not write) |
create_from_template |
template_name (optional) |
Discover templates (if needed), read a template, gather user values, and write a new note |
related |
path |
Find related notes via search and suggest cross-references as markdown links |
compare |
path1, path2 |
Read two documents and produce a side-by-side comparison |
propose-links |
scope (optional), per_note_limit (optional) |
Scan a candidate set of notes (a folder, recent, or all), propose links between semantically close notes that aren't already connected, and write them on confirmation |
Write prompts (research, discuss, create_from_template, propose-links) are only available when MARKDOWN_VAULT_MCP_READ_ONLY=false.
Templates are regular markdown files. If placeholder template text pollutes search results, add your templates folder to MARKDOWN_VAULT_MCP_EXCLUDE (such as _templates/**).
Mount a directory of .md prompt files to override or extend the built-in prompts. Set MARKDOWN_VAULT_MCP_PROMPTS_FOLDER to the path. Each file's frontmatter defines description, arguments (a list of objects, each with name, description, and required fields), and optional tags. A user prompt with the same name as a built-in replaces it.
For a complete example, including Zettelkasten capture, development, and review prompts, see the Zettelkasten guide. For an alternative action-oriented workflow (Projects, Areas, Resources, Archive with triage, kickoff, and weekly review prompts), see the PARA guide.
The server ships four browser-based views that MCP clients supporting the MCP Apps protocol can render inline or in fullscreen. They are delivered as a single HTML resource at ui://markdown_vault_mcp/app.html and registered using visibility="app" so they appear only in supporting clients and do not clutter the standard tool list. See the MCP Apps guide for details.
| View | Description |
|---|---|
| Context Card | Displays a note dossier (backlinks, outlinks, similar notes, tags) for the note currently in focus |
| Graph Explorer | Interactive force-directed link graph of the vault, powered by vis-network |
| Vault Browser | Searchable, filterable file tree for navigating the vault without issuing tool calls |
| Note Preview | Full-width markdown preview with a Contents popover, collapsible frontmatter properties and tags, copy-markdown / copy-vault-link controls, and a "Send to Claude" button |
The two primary tools exposed to MCP Apps clients are:
| Tool | Description |
|---|---|
browse_vault |
Returns the vault tree structure for the Vault Browser view |
show_context |
Returns the full context dossier for a given note path (used by the Context Card view) |
Domain configuration: MCP Apps iframes are sandboxed to a specific Claude app domain. The domain is auto-computed from MARKDOWN_VAULT_MCP_BASE_URL. Override with MARKDOWN_VAULT_MCP_APP_DOMAIN if your deployment is hosted on a custom domain or behind a proxy that changes the apparent hostname.
Vendored dependencies (JavaScript libraries bundled at build time, no runtime CDN): vis-network (graph rendering), marked.js (markdown rendering), DOMPurify (XSS sanitization), ext-apps SDK (MCP Apps lifecycle). The one runtime network dependency is web fonts (Newsreader, Public Sans, IBM Plex Mono), loaded from Google Fonts with system-font fallbacks.
create_download_link and create_upload_link mint short-lived capability URLs so vault files can move to a browser or another service without inflating the LLM context window. The token embedded in the URL is the only credential; no Authorization header is required on the /transfer/{token} route.
# Download a vault file
create_download_link(path="reports/q1.pdf", ttl_seconds=600)
# → {"url": "https://mcp.example.com/transfer/<token>", ...}
curl "https://mcp.example.com/transfer/<token>" -o q1.pdf
# Upload a file to the vault
create_upload_link(path="assets/new-diagram.png")
# → {"url": "https://mcp.example.com/transfer/<token>", ...}
curl -X POST --data-binary @new-diagram.png "https://mcp.example.com/transfer/<token>"
Each token is consumed on its first successful use. A failed or interrupted transfer does not burn the token; retry is permitted until the TTL expires.
Requirements: HTTP or SSE transport; MARKDOWN_VAULT_MCP_BASE_URL set. See the transfer links guide for the full walkthrough and security model.
Beyond Markdown notes, the server can read, write, delete, rename, and list non-markdown files (PDFs, images, spreadsheets, etc.). All existing tools are overloaded; there are no new tool names.
Path dispatch is extension-based: a path ending in .md is treated as a note; any other path is treated as an attachment if the extension is in the allowlist. The kind field on returned objects distinguishes the two: "note" or "attachment".
read returns base64-encoded content for binary attachments:
{
"path": "assets/diagram.pdf",
"mime_type": "application/pdf",
"size_bytes": 12345,
"content_base64": "<base64 string>",
"modified_at": 1741564800.0
}
write accepts a content_base64 parameter for binary content:
{ "path": "assets/diagram.pdf", "content_base64": "<base64 string>" }
list_documents with include_attachments=true returns both notes and attachments:
[
{ "path": "notes/intro.md", "kind": "note", "title": "Intro", "folder": "notes", "frontmatter": {}, "modified_at": 1741564800.0 },
{ "path": "assets/diagram.pdf", "kind": "attachment", "folder": "assets", "mime_type": "application/pdf", "size_bytes": 12345, "modified_at": 1741564800.0 }
]
pdf, docx, xlsx, pptx, odt, ods, odp, png, jpg, jpeg, gif, webp, svg, bmp, tiff, zip, tar, gz, mp3, mp4, wav, ogg, txt, csv, tsv, json, yaml, toml, xml, html, css, js, ts
Override with MARKDOWN_VAULT_MCP_ATTACHMENT_EXTENSIONS. Use * to allow all non-.md files.
Hidden directories: Attachments inside hidden directories (
.git/,.obsidian/,.markdown_vault_mcp/, etc.) are never listed, regardless of extension settings.MARKDOWN_VAULT_MCP_EXCLUDEpatterns are also applied to attachments.
The server supports four auth modes:
MARKDOWN_VAULT_MCP_BEARER_TOKEN to a secret stringOIDC_CONFIG_URL, OIDC_CLIENT_ID, OIDC_CLIENT_SECRET, and BASE_URLAuth requires --transport http (or sse). It has no effect with --transport stdio.
For setup instructions, troubleshooting, and provider-specific guides, see the Authentication guide.
git clone https://github.com/pvliesdonk/markdown-vault-mcp.git
cd markdown-vault-mcp
uv sync --all-extras --all-groups
# Run tests
uv run python -m pytest tests/ -x -q
# Lint and format
uv run ruff check src/ tests/
uv run ruff format src/ tests/
# Type check
uv run mypy src/ tests/
CI workflows reference three repository secrets. Configure them via Settings → Secrets and variables → Actions or with gh secret set:
| Secret | Used by | How to generate |
|---|---|---|
RELEASE_TOKEN |
release.yml, copier-update.yml, renovate.yml, bootstrap.yml |
Fine-grained PAT at https://github.com/settings/personal-access-tokens/new with contents: write, pull_requests: write, and administration: write (bootstrap sets branch protection + auto-merge). Scoped to this repo. |
CODECOV_TOKEN |
ci.yml |
Sign in at https://codecov.io with GitHub, add the repo, then copy the upload token from the repo settings page. |
CLAUDE_CODE_OAUTH_TOKEN |
claude.yml, claude-code-review.yml |
Run claude setup-token locally and paste the result. |
GITHUB_TOKEN is auto-provided; no action needed.
Dependency updates are handled by Renovate (
renovate.yml), which reusesRELEASE_TOKEN. It maintainsuv.lockand auto-merges patch/minor bumps once theCI Successcheck is green;bootstrap.ymlenables auto-merge and branch protection on first push. GitHub Actions are updated in the copier template and arrive viacopier update, not per-repo.
uv sync creates .venv/bin/* scripts with absolute shebangs pointing at the venv Python. If you move the repo (mv /old/path /new/path), uv run pytest fails with ModuleNotFoundError because the stale shebang resolves to a different interpreter than the venv's site-packages.
Fix:
rm -rf .venv
uv sync --all-extras --all-groups
uv run python -m pytest also works as a one-shot workaround.
uv.lock refresh after copier updateWhen copier update introduces new dependencies, CI runs uv sync --frozen which fails against a stale lockfile. Run uv lock locally and commit the refreshed uv.lock alongside accepting the copier-update PR.
Package root minimized (issue #903): import from submodules, not the package root.
The markdown_vault_mcp root package no longer re-exports the public API. Update
library imports to their submodules:
# Before: from markdown_vault_mcp import Vault, ProjectConfig
# After: from markdown_vault_mcp.vault import Vault
# from markdown_vault_mcp.config import ProjectConfig
Types such as GroupedResult come from markdown_vault_mcp.types. Only
__version__ remains importable from the root.
v2.0.0 (issue #469): search, get_similar, and get_context.similar now return grouped results.
Each file appears once with a sections list; the flat content, heading, and score
fields have moved inside each SectionHit. Library consumers must update iteration:
# Before: result.content, result.heading
# After: result.sections[0].content, result.sections[0].heading
MARKDOWN_VAULT_MCP_CHUNKS_PER_FILE replaces MARKDOWN_VAULT_MCP_CHUNKS_PER_DOC.
SimilarItem is removed; use GroupedResult (from markdown_vault_mcp.types).
MARKDOWN_VAULT_MCP_MAX_ATTACHMENT_SIZE_MB default lowered from 10 MB
to 1 MB. Most LLM contexts can't survive a 10 MB base64-encoded
attachment; the old default was a silent context-blow-up. If you have
non-LLM consumers (scripts, CI) that need the old behaviour, set
MARKDOWN_VAULT_MCP_MAX_ATTACHMENT_SIZE_MB=10 explicitly.
MARKDOWN_VAULT_MCP_MAX_NOTE_READ_BYTES is a new env var (default
256 KB). Whole-document .md reads above this raise ValueError.
Partial reads via read(path, section=heading) bypass the cap.
MIT
Please log in to share your review and rating for this MCP.
Explore related MCPs that share similar capabilities and solve comparable challenges
by exa-labs
Provides real-time web search capabilities to AI assistants via a Model Context Protocol server, enabling safe and controlled access to the Exa AI Search API.
by perplexityai
Enables Claude and other MCP‑compatible applications to perform real‑time web searches through the Perplexity (Sonar) API without leaving the MCP ecosystem.
by Aas-ee
Provides multi-engine web search and article/content retrieval through an MCP server, command‑line interface, and a reusable local daemon, all without requiring API keys or authentication.
by MicrosoftDocs
Provides semantic search and fetch capabilities for Microsoft official documentation, returning content in markdown format via a lightweight streamable HTTP transport for AI agents and development tools.
by elastic
Enables natural‑language interaction with Elasticsearch indices via the Model Context Protocol, exposing tools for listing indices, fetching mappings, performing searches, running ES|QL queries, and retrieving shard information.
by graphlit
Enables integration between MCP clients and the Graphlit platform, providing ingestion, extraction, retrieval, and RAG capabilities across a wide range of data sources and connectors.
by ihor-sokoliuk
Provides web search capabilities via the SearXNG API, exposing them through an MCP server for seamless integration with AI agents and tools.
by mamertofabian
Fast cross‑platform file searching leveraging the Everything SDK on Windows, Spotlight on macOS, and locate/plocate on Linux.
by spences10
Provides unified access to multiple search engines, AI response tools, and content processing services through a single Model Context Protocol server.