by thomas-villani
Convert PDFs, Word, PowerPoint, HTML, email and over 40 other formats into clean, LLM‑ready Markdown and back, with a Python API and powerful CLI.
All2md provides a Python library and command‑line interface that parses a huge variety of document types (PDF, DOCX, PPTX, HTML, email, spreadsheets, ebooks, source code, archives, etc.) into a structured, Markdown representation optimized for large language model (LLM) ingestion. The conversion is bidirectional, allowing Markdown to be rendered back into rich formats such as DOCX, PDF, PPTX, HTML, EPUB, and custom text formats via Jinja2 templates.
pip install "all2md[pdf]") and run commands like all2md report.pdf > report.md or all2md notes.md --out notes.docx. Batch conversion, recursive directory processing, grepping, chunking, previewing, and diffing are all supported.to_markdown, convert, chunk, to_ast, from_ast, etc. Example: markdown = to_markdown("document.pdf") or chunks = all2md.chunk("paper.pdf", strategy="semantic", max_tokens=512). Typed options objects allow fine‑grained control (OCR, page selection, renderer flavor, etc.).How does All2md differ from Pandoc? All2md is Python‑native, focuses on programmatic use, LLM integration, and extensibility, while Pandoc is a broader Haskell‑based publishing tool.
Can I convert Markdown back to Word or PDF? Yes – use all2md input.md --out output.docx or the convert API function.
What is the recommended Markdown flavor for LLMs? GitHub Flavored Markdown (gfm) provides consistent structure and is well‑understood by most models.
Does it handle scanned PDFs? Yes, with OCR extras (all2md[pdf,ocr] or all2md[pdf,ocr-easyocr]). Enable with --pdf-ocr-enabled or the corresponding options object.
Can I output formats other than Markdown? Custom Jinja2 templates allow rendering the AST to any text‑based format such as DocBook XML, YAML, or ANSI.
How do I add a new file format? Implement a parser class, define ConverterMetadata, and register it via the all2md.converters entry point in pyproject.toml.
How to speed up large batch conversions? Use parallel workers (-p N), the --watch mode for incremental processing, and enable the on‑disk cache (--cache).
Convert PDFs, Office files, HTML, emails, spreadsheets, and 40+ other formats into clean, LLM-ready Markdown — and back again.
all2md is a Python library and command-line tool for turning many document formats into structured, LLM-friendly Markdown — and converting Markdown back into rich formats like DOCX, PDF, and HTML. Built on an AST-based pipeline, it's designed for RAG ingestion, LLM preprocessing, batch automation, and embedding document conversion directly into Python applications.
📦 PyPI · 📖 Documentation · 💡 Examples
# Install with PDF support (add more extras as you need them)
pip install "all2md[pdf]"
# Convert any document to Markdown (prints to stdout)
all2md report.pdf > report.md
# Go the other way — Markdown back to a rich format
all2md notes.md --out notes.docx
In Python:
from all2md import to_markdown
markdown = to_markdown("report.pdf")
That's it. For more formats, install only the extras you need — all2md[docx,html,xlsx] — or all2md[all] for everything.
# Convert a PDF to Markdown for RAG / LLM ingestion
all2md paper.pdf > paper.md
# Batch-convert a directory (recursively) into a folder of Markdown
all2md ./docs --recursive --output-dir ./markdown
# Grep across mixed document types like they were plain text
all2md grep "revenue" reports/*.pdf
# Chunk a document for a RAG pipeline (JSONL with section + page provenance)
all2md chunk handbook.pdf --strategy semantic --max-tokens 512 --overlap 64
# Preview any document in your browser
all2md view proposal.docx
# Turn Markdown (e.g. an LLM's output) back into DOCX, PDF, or PPTX
all2md answer.md --out answer.docx
Every file command supports stdin/stdout via -, so you can pipe and chain:
curl -s https://example.com/doc.pdf | all2md - | grep "important"
Reach for all2md when you want a Python-first, automation-friendly document workflow with first-class LLM integration. Reach for Pandoc when you need maximum publishing breadth or advanced scholarly output (citations, bibliographies). They complement each other well.
A PDF research paper in, structured Markdown out — headings, prose, and tables preserved:
# Efficient Retrieval Methods
## Abstract
We study retrieval-augmented generation across a range of...
## 1 Introduction
Retrieval-augmented generation (RAG) combines a retriever with...
| Model | Accuracy | Latency |
|-------|---------:|--------:|
| A | 91.2% | 40 ms |
| B | 93.8% | 65 ms |
Tables, multi-column layouts, and scanned pages (via OCR) are handled by the advanced PDF parser. See all2md report to score how much to trust any given conversion.
Word documents come out as Word shows them: tracked changes resolved by policy (--docx-revisions accept|reject|mark), comment threads with their anchors and replies, footnote and endnote bodies, field results, content controls, merged cells, and the list numbers and labels Word actually prints.
all2md uses a modular system — dependencies are only required for the formats you actually process.
Run all2md list-formats to see everything on your install, or browse the full formats matrix.
| Format | File Extensions | Input (Parse) | Output (Render) | Dependencies Extra |
|---|---|---|---|---|
.pdf |
✅ | ✅ | pdf, pdf_render |
|
| Word Document | .docx |
✅ | ✅ | docx |
| PowerPoint Presentation | .pptx |
✅ | ✅ | pptx |
| HTML | .html, .htm |
✅ | ✅ | html |
| MHTML Web Archive | .mhtml, .mht |
✅ | (N/A) | html |
| Email Message | .eml |
✅ | (N/A) | (built-in) |
| MBOX Mailbox Archive | .mbox, .mbx |
✅ | (N/A) | (built-in) |
| Outlook Message/Archive | .msg, .pst, .ost |
✅ | (N/A) | outlook |
| Jupyter Notebook | .ipynb |
✅ | ✅ | (built-in) |
| EPUB E-book | .epub |
✅ | ✅ | epub |
| FictionBook 2.0 (FB2) | .fb2 |
✅ | (N/A) | fb2 |
| CHM (Compiled HTML Help) | .chm |
✅ | (N/A) | chm |
| OpenDocument Text | .odt |
✅ | ✅ | odf |
| OpenDocument Presentation | .odp |
✅ | ✅ | odf |
| OpenDocument Spreadsheet | .ods |
✅ | (N/A) | odf |
| Excel Spreadsheet | .xlsx |
✅ | (N/A) | xlsx |
| CSV / TSV | .csv, .tsv |
✅ | ✅ | (built-in) |
| Rich Text Format | .rtf |
✅ | ✅ | rtf |
| LaTeX | .tex, .latex |
✅ | ✅ | latex |
| AsciiDoc | .adoc, .asciidoc, .asc |
✅ | ✅ | (built-in) |
| reStructuredText | .rst |
✅ | ✅ | rst |
| Org-Mode | .org |
✅ | ✅ | org |
| MediaWiki | .wiki, .mw |
✅ | ✅ | wiki |
| Textile | .textile |
✅ | ✅ | (built-in) |
| BBCode | .bbcode, .bb |
✅ | (N/A) | (built-in) |
| DokuWiki | .doku, .dokuwiki |
✅ | ✅ | (built-in) |
| Evernote Export | .enex |
✅ | (N/A) | enex |
| Safari Web Archive | .webarchive |
✅ | (N/A) | html |
| JSON | .json |
✅ | ✅ | (built-in) |
| YAML | .yaml, .yml |
✅ | ✅ | (built-in) |
| TOML | .toml |
✅ | ✅ | (built-in) |
| INI / Config | .ini, .cfg, .conf |
✅ | ✅ | (built-in) |
| OpenAPI/Swagger | .yaml, .yml, .json |
✅ | (N/A) | openapi |
| Plain Text | .txt, .text |
✅ | ✅ | (built-in) |
| Source Code | nearly 200 extensions (.py, .js, etc.) |
✅ | (N/A) | (built-in) |
| Archive Formats | .tar, .tgz, .7z, .rar, etc. |
✅ | (N/A) | (built-in) |
| ZIP Archive | .zip |
✅ | (N/A) | (built-in) |
| Jinja2 Templates (Custom) | User-defined (.jinja2, .j2) |
❌ | ✅ | jinja2 |
💡 Custom output formats: render to any text-based format using Jinja2 templates, no Python required. See the Template Guide and examples/templates/.
The core library has no dependencies — install support for formats as you need them.
CLI (system-wide, no Python setup to manage):
uv tool install "all2md[all]"
Python library:
pip install "all2md[pdf,docx,html]"
Minimal (core only):
pip install all2md
Check what format support you have installed:
all2md check-deps
One-click install (no Python setup required). The scripts set up uv (installing it first if needed) and install the all2md CLI globally.
macOS / Linux (bash or zsh):
curl -LsSf https://raw.githubusercontent.com/thomas-villani/all2md/main/scripts/install.sh | sh
Windows (PowerShell):
powershell -ExecutionPolicy Bypass -c "irm https://raw.githubusercontent.com/thomas-villani/all2md/main/scripts/install.ps1 | iex"
Both scripts install the all extra by default. To slim it down, download the script and pass a comma-separated extras list — sh install.sh pdf,docx,html or .\install.ps1 -Extras pdf,docx,html (use none for a base-only install). The scripts are also attached to each GitHub release.
Extras for specific needs:
# Spreadsheets and ODF documents
pip install "all2md[xlsx,odf]"
# PDF with OCR for scanned documents (Tesseract engine; needs the system binary)
pip install "all2md[pdf,ocr]"
# ...or the binary-free EasyOCR engine (downloads models on first use)
pip install "all2md[pdf,ocr-easyocr]"
# PDF with GNN-based semantic layout analysis
pip install "all2md[pdf_layout]"
# Outlook MSG files
pip install "all2md[outlook]"
# Note: PST/OST support requires an extra manual step: pip install libpff-python
# Everything
pip install "all2md[all]"
The essentials:
all2md document.pdf # convert to Markdown on stdout
all2md report.docx --out report.md # write to a file
all2md notes.md --out notes.docx # Markdown → rich format (bidirectional)
all2md ./docs -r --output-dir ./out # recursively batch-convert a directory
all2md document.pdf --rich # render in the terminal (fancy `cat`)
all2md view document.pdf --theme docs # HTML preview in the browser
all2md grep "search term" documents/*.pdf # grep through any document format
# View & edit
all2md doc.pdf --rich # rich terminal rendering (rcat = shorthand)
all2md view document.pdf # instant HTML preview in the browser
all2md serve ./docs --recursive # serve a directory over HTTP with live preview
all2md edit notes.md # browser-based Markdown/WYSIWYG editor, saves back
# Extract sections by heading
all2md doc.pdf --extract "Introduction"
all2md view report.docx --extract "Q3 Results"
# Grep and search
all2md grep -i "case insensitive" report.docx
all2md search "machine learning" ./research_papers/
all2md search "project timeline" --semantic ./docs/
# Chunk documents for RAG/LLM pipelines (JSONL with section + page provenance)
all2md chunk report.pdf --strategy semantic --max-tokens 512 --overlap 64
# Conversion quality: score, round-trip, and auto-tune
all2md report scan.pdf # reference-free quality "card"
all2md report inbox/*.docx --fail-under 80 # CI gate
all2md roundtrip notes.md --via docx # convert → parse back → score fidelity
all2md optimize scanned.pdf --sample-pages 5 # auto-tune settings for a hard document
# Diff any two documents (any format), like Unix diff
echo "<p>Version 1</p>" | all2md diff - version2.html
# Package a paper for ArXiv submission
all2md arxiv paper.md -o submission.tar.gz --bib references.bib
# Multi-file, parallel, and watch mode
all2md ./large_docs -r --output-dir ./output -p 4
all2md ./watched_folder -r --output-dir ./output --watch
# Apply AST transforms from the CLI
all2md report.docx -t remove-images
all2md chapter.docx -t "heading-offset --offset 1"
all2md list-transforms
# Static site generation (Hugo, Jekyll, MkDocs, Zola, Eleventy)
all2md generate-site ./content --output-dir ./site --generator hugo --scaffold
# Format-specific options — every option is a CLI flag; run --help to see them
all2md report.pdf --pdf-pages "1-3,5"
all2md scanned.pdf --pdf-ocr-enabled --pdf-ocr-mode auto --pdf-ocr-languages eng
all2md document.docx --attachment-mode save --attachment-output-dir ./images
# Discovery & config
all2md list-formats
all2md config generate > all2md.toml
all2md config validate all2md.toml
Speed up repeat runs with the opt-in on-disk cache (--cache, or export ALL2MD_CACHE=1), which reuses parsed documents across grep, search, chunk, view, report, roundtrip, and optimize.
The to_markdown() function is the easiest way to get started; convert() handles conversions between any two formats.
from all2md import to_markdown, convert
# Convert a file to Markdown
markdown = to_markdown("document.pdf")
# Fine-tune with typed options or plain keyword arguments
markdown = to_markdown("report.pdf", pages="1-3,5", flavor="gfm")
# Bidirectional conversion between any two supported formats
convert("input.md", "output.docx", target_format="docx")
convert("page.html", "page.pdf", target_format="pdf")
Chunking for RAG — convert and split in one call, keeping provenance most chunkers throw away:
import all2md
chunks = all2md.chunk("report.pdf", strategy="semantic", max_tokens=512, overlap=64)
for c in chunks:
print(c.chunk_id, c.section_heading, c.page, c.token_count)
record = c.to_dict() # flat dict — the same object emitted as JSONL by the CLI
Typed options objects give you type safety and clarity:
from all2md import to_markdown, PdfOptions, MarkdownRendererOptions
pdf_opts = PdfOptions(pages="1-3,5", attachment_mode="base64")
md_opts = MarkdownRendererOptions(flavor="gfm", emphasis_symbol="_")
markdown = to_markdown("report.pdf", parser_options=pdf_opts, renderer_options=md_opts)
# Scanned PDF with OCR (engine="tesseract" default, or "easyocr" for binary-free)
from all2md.options.common import OCROptions
ocr_opts = OCROptions(enabled=True, mode="auto", engine="tesseract", languages="eng", dpi=300)
markdown = to_markdown("scanned.pdf", parser_options=PdfOptions(ocr=ocr_opts))
Working with the AST for advanced processing:
from all2md import to_ast, from_ast
from all2md.ast import Heading, Text
doc = to_ast("document.pdf") # parse to AST
doc.children.insert(0, Heading(level=1, content=[Text(content="New Title")]))
markdown = from_ast(doc, target_format="markdown") # render back out
Transform pipelines for systematic modification:
from all2md import to_ast
from all2md.transforms import render, HeadingOffsetTransform, RemoveImagesTransform
doc = to_ast("report.docx")
markdown = render(doc, transforms=[
RemoveImagesTransform(), # an instance
"add-heading-ids", # a registered name
HeadingOffsetTransform(offset=1), # an instance with parameters
])
Real BPE token counting for chunking uses tiktoken (pip install all2md[chunk]); count-only strategies fall back to a whitespace approximation. See the API documentation for the full reference and more examples under examples/python/.
all2md is built to sit inside LLM and agent workflows.
all2md chunk (and all2md.chunk()) split any document into retrieval-ready chunks, each carrying its section heading/level and source page span so answers can cite where they came from. 11 strategies (semantic/heading/section/token/sentence/paragraph/word/line/char/code/auto); keep tables and code blocks whole; strip noisy elements.all2md install-skills, or get the same guidance without installing anything via all2md llm-help [topic].pip install "all2md[mcp]"
all2md-mcp --temp --enable-from-md
Tools: read_document_as_markdown, save_document_from_markdown, edit_document, plus three read-only query tools enabled by default — search_documents (grep + keyword/BM25 across a corpus), diff_documents, and get_document_outline.
One-click install (Claude Desktop). Install the prebuilt MCPB bundle — no manual config or separate Python install required (the bundle pulls in all2md via uv on first run):
all2md.mcpb from the latest release.all2md.mcpb onto the Extensions pane (or use Install Extension to browse for it).Requires a Claude Desktop build with MCPB extension support (late-2025 or newer). To rebuild the bundle yourself, see
mcpb/README.md.
Manual configuration (developers / other MCP clients) — add to claude_desktop_config.json:
{
"mcpServers": {
"all2md": {
"command": "all2md-mcp",
"args": ["--temp", "--enable-from-md"]
}
}
}
Using uvx (which does not require system installation of all2md):
{
"mcpServers": {
"all2md": {
"command": "uvx",
"args": ["--from", "all2md[all]", "all2md-mcp", "--temp", "--enable-from-md"]
}
}
}
See the MCP documentation and Agent Skills documentation for full details.
Most document tooling in CI answers "did it run?". all2md ships a GitHub Action that answers "is the output still as good as it was?" — it scores every matched document and fails the build when fidelity degrades:
- uses: thomas-villani/all2md@v1.15.1
with:
paths: docs/**/*.md
roundtrip-fail-under: 97
Measure your real floor before picking a threshold — all2md roundtrip docs/*.md --fail-under 1 prints it. Documents that convert well score 99–100, so a threshold that sounds strict (80, say) can have twenty points of dead headroom and never fire. The action warns you when that happens.
It also refuses to pass quietly: no matching files, no threshold set, or a document that cannot be converted at all are all failures rather than silent greens. The same gate works without the Action, in any CI system — see the full documentation.
Built on an AST-based pipeline (parse → transform → render), all2md offers capabilities that direct format-to-format converters can't:
pdf_layout extra), header/footer removal, and OCR for scanned pages (Tesseract or binary-free EasyOCR) — powered by PyMuPDF, measured against external ground truth (see below).all2md report gives a reference-free confidence "quality card" for any document (usable as a CI gate); all2md roundtrip scores how much structure survives a convert → parse-back round trip; all2md optimize auto-tunes converter settings for a difficult document.diff command that works like Unix diff but across any document formats, with text-based symmetric comparison.all2md.converters entry point) and transforms (all2md.transforms entry point). See examples/plugins/.all2md's conversion quality is measured, not asserted — against three independent ground truths, each published beside a control that shows what the measurement looks like when it should fail:
benchmarks/pmc/): 66 publisher PDFs from PubMed Central scored against the JATS XML deposited beside them. Text recall, structural recovery of headings and tables, and an invented-text rate — with the corpus pinned by committed SHA-256 digests.benchmarks/omnidocbench/): 981 raster pages against human annotation, exercising the OCR path the born-digital corpus never touches.benchmarks/roundtrip/): Markdown → format → Markdown must survive at fidelity 100 for the repository's own docs, gated on every pull request.Current figures, their controls, and — just as important — what each lane structurally cannot see are documented in Conversion Fidelity.
How does it compare to other converters? A fourth lane (benchmarks/comparison/) scores pymupdf4llm and Docling with the same instruments, on a corpus held out from the born-digital lane, with every tool's output re-parsed through one normalization path. In the reading of 2026-08-29 — taken on a freshly drawn, sealed holdout against a corrected ground truth — all2md has by far the lowest invented-text rate (0.55%, vs 6.26% for pymupdf4llm, whose defaults now auto-OCR born-digital pages, and 2.05% for Docling), is the fastest of the three, and leads recall of what is attainable (97.3% to 97.1% and 96.2%), while Docling leads table cell preservation by 6.2 points (79.3% to 73.1%). The first reading on that holdout had put the table gap at 16.7 points; 7.1 of those were a ground-truth artifact that hid a whole class of table from every tool, and 4.1 were closed by row-grouping fixes since. What remains is diagnosed in the lane as a row-count difference — a learned table model against geometry rules — rather than a rule all2md is missing. Full results, ground rules, and the caveats that bound them are in that lane's README and dated results-*.json snapshots.
How is all2md different from Pandoc? all2md is Python-native, with a focus on programmatic use, LLM integration, and extensibility. Pandoc is more comprehensive for scholarly documents but is Haskell-based and CLI-focused. Use all2md for Python projects and AI workflows; use Pandoc for academic publishing — they complement each other well.
Can I convert back from Markdown to Word/PDF?
Yes — all2md is bidirectional. Use convert("input.md", "output.docx", target_format="docx"), or the CLI: all2md input.md --out output.pdf.
What's the best format for feeding documents to LLMs?
Markdown with the gfm (GitHub Flavored Markdown) flavor — structured, consistent, and well-understood by LLMs. Use to_markdown(file, flavor="gfm").
Does all2md work with scanned PDFs?
Yes. Install OCR support (pip install "all2md[pdf,ocr]") and use --pdf-ocr-enabled (or OCROptions(enabled=True)). The default Tesseract engine needs the Tesseract binary; the binary-free EasyOCR engine is available via all2md[pdf,ocr-easyocr] and --pdf-ocr-engine easyocr.
Can I customize the output beyond Markdown? Yes — use Jinja2 templates to render any text-based format. See examples/templates/ for DocBook XML, YAML, ANSI terminal output, and more.
How do I add support for a new file format?
Create a parser class, define a ConverterMetadata object, and register it via the all2md.converters entry point in your pyproject.toml. See examples/plugins/ for a complete example.
How do I handle large document batches efficiently?
Use parallel processing: all2md ./docs -r --output-dir ./output -p 8, or --watch for incremental processing as files arrive.
Contributions are welcome — bug reports, feature requests, documentation improvements, and code. Ways to help: report bugs, improve docs, add support for new formats via the plugin system, create new AST transforms, or fix bugs in existing converters.
For contributors evaluating parser changes, all2md ships benchmark harnesses: benchmarks/pmc/ (born-digital PDF fidelity against publisher JATS ground truth), benchmarks/omnidocbench/ (scanned pages against human annotation), benchmarks/roundtrip/ (Markdown → AST → Markdown fidelity, the CI gate), benchmarks/corpus/ (conversion timing across public corpora), and benchmarks/comparison/ (third-party converters scored with the same instruments on a held-out corpus). See Conversion Fidelity and the Performance Tuning docs for details.
See CONTRIBUTING.md for development setup and guidelines.
This project is licensed under the MIT License. See the LICENSE file for details.
Please log in to share your review and rating for this MCP.
Explore related MCPs that share similar capabilities and solve comparable challenges
by headroomlabs-ai
Compress tool outputs, logs, files, RAG chunks, and conversation history before they reach the LLM, keeping answers identical while saving up to 95% of tokens for JSON payloads.
by modelcontextprotocol
A Model Context Protocol server for Git repository interaction and automation.
by zed-industries
A high‑performance, multiplayer code editor designed for speed and collaboration.
by modelcontextprotocol
Model Context Protocol Servers
by modelcontextprotocol
A Model Context Protocol server that provides time and timezone conversion capabilities.
by cline
An autonomous coding assistant that can create and edit files, execute terminal commands, and interact with a browser directly from your IDE, operating step‑by‑step with explicit user permission.
by upstash
Provides up-to-date, version‑specific library documentation and code examples directly inside LLM prompts, eliminating outdated information and hallucinated APIs.
by daytonaio
Provides a secure, elastic infrastructure that creates isolated sandboxes for running AI‑generated code with sub‑90 ms startup, unlimited persistence, and OCI/Docker compatibility.
by continuedev
Enables faster shipping of code by integrating continuous AI agents across IDEs, terminals, and CI pipelines, offering chat, edit, autocomplete, and customizable agent workflows.