A quiet little Python tool just shot to the top of GitHub's trending charts — 100,000+ stars. It's called MarkItDown.
— @mdancho84 September 6, 2026
Note the engagement datapoint honestly: the viral post says “100,000+”; the live repo shows ~180k (checked September 7, 2026 — the star count moves fast, date-stamp any citation).
What MarkItDown Is
Microsoft's own framing: “a lightweight Python utility for converting various files to Markdown for use with LLMs and related text analysis pipelines.” The design goal is LLM-consumable structure, not visual fidelity — headings, lists, tables, and links survive; pixel-perfect layout does not. That is a feature for RAG and agent pipelines (tokens and structure matter more than typography) and a limitation for human-facing document reproduction.
Format Support
| Category | Supported | Notes |
|---|---|---|
| Documents | PDF, DOCX, PPTX, XLSX/XLS | Text-layer extraction; structure preserved (headings, lists, tables). |
| Images | JPEG, PNG (EXIF + OCR) | OCR applies to image files — not to images embedded inside PDFs. |
| Audio | WAV, MP3 (EXIF + transcription) | Basic transcription built in; advanced audio via Azure. |
| Web / data | HTML, CSV, JSON, XML, RSS, Wikipedia, YouTube URLs | YouTube pulls transcripts. |
| Containers / ebooks | ZIP (iterates contents), EPub, Outlook messages, Jupyter notebooks | — |
CLI and Python Usage
Install with the works: pip install 'markitdown[all]' (Python ≥3.10) — or a slim subset like pip install 'markitdown[pdf, docx, pptx]'. CLI one-liners: markitdown file.pdf > out.md, cat file.pdf | markitdown, or markitdown file.pdf -o out.md. Python: from markitdown import MarkItDown; md = MarkItDown(); print(md.convert("test.xlsx").text_content). Narrower entry points (convert_local, convert_stream, convert_response) constrain what untrusted input can reach — prefer them when handling untrusted files. Docker works too: docker run --rm -i markitdown:latest < file.pdf > out.md.
Power features need Azure or an LLM client: Document Intelligence for high-fidelity PDFs (-d -e "<endpoint>"), Azure Content Understanding for richer media analysis, and GPT-4o-class captions for images inside PPTX (MarkItDown(llm_client=OpenAI(), llm_model="gpt-4o")). Those paths are billable Azure/OpenAI usage — the core converter is not.
The Scanned-PDF Limit and Other Gotchas
The gotcha everyone hits: scanned or image-only PDFs return near-empty output — the built-in PDF path extracts the text layer (pdfminer.six), and its OCR applies to standalone image files, not images embedded inside PDFs. For scans, route through Azure Document Intelligence/Content Understanding or the community markitdown-ocr plugin (LLM Vision). Other gotchas: plugins are opt-in (enable_plugins/flag) because they execute third-party code; the tool runs with your process privileges, so sanitize untrusted input and restrict paths and URI schemes; and MCP exposure (next section) deliberately does not include the billable or plugin features.
The MCP Server
markitdown-mcp exposes exactly one tool — convert_to_markdown(uri) — over stdio, SSE, or Streamable HTTP. That minimalism is deliberate: your agent gets document conversion without access to LLM-caption billing, Azure endpoints, or plugin execution. For agent pipelines, that is the right default; for richer extraction, call the Python API in your own tool layer.
Use It If
Use MarkItDown if you feed documents to LLMs — RAG ingestion, agent document-reading, dataset prep — and want one MIT-licensed, zero-cost dependency for the common formats. Pair it or skip it if your corpus is scanned (add DocIntel or the OCR plugin), you need visual-fidelity conversion for humans, or you need benchmarked accuracy guarantees — independent accuracy comparisons against Docling and peers are not yet published, so validate on your own documents.
A RAG Ingestion Recipe
The pattern that fits most document pipelines — convert narrow, route scans elsewhere, keep the MCP surface minimal:
from markitdown import MarkItDown
md = MarkItDown(enable_plugins=False) # never run third-party plugins on untrusted input
def to_markdown(path: str) -> str:
if path.lower().endswith((".pdf", ".docx", ".pptx", ".xlsx")):
result = md.convert_local(path) # narrow API = process-privilege safety
if len(result.text_content.strip()) < 40:
return route_to_ocr(path) # scanned file: Azure DocIntel / markitdown-ocr
return result.text_content
return md.convert_local(path).text_contentWhy this shape: enable_plugins=False keeps untrusted input away from third-party code; the near-empty check catches scanned PDFs (the built-in extractor reads text layers only); and convert_local restricts what a malicious file can reach. Agent-facing exposure stays minimal via markitdown-mcp's single convert_to_markdown(uri) tool.
FAQ
Why is my scanned PDF coming out empty?
Built-in PDF conversion reads the text layer only. For scans, use Azure Document Intelligence (-d -e endpoint), Azure Content Understanding, or the markitdown-ocr plugin.
Does MarkItDown preserve tables?
It preserves table structure as Markdown tables where the source format exposes it (Office, HTML). Complex nested or merged-cell layouts flatten — the goal is LLM-consumable structure, not reproduction.
Is it free?
The core converter is MIT-licensed and free. Billable paths are optional: Azure Document Intelligence, Content Understanding, and LLM-generated captions.
MarkItDown vs Docling?
No rigorous independent accuracy benchmark is published as of September 2026. MarkItDown's advantages are Microsoft backing, the MIT license, breadth of formats, and the minimal MCP server. Evaluate both on your own corpus — converter accuracy is document-distribution-dependent.
Sources
- GitHub — microsoft/markitdown (README, releases, LICENSE)
- MarkItDown docs — introduction and quickstart
Get the latest on AI, LLMs & developer tools
New MCP servers, model updates, and guides like this one — delivered weekly.