Cost & tokens

MarkItDown: Convert PDFs to Markdown Before Claude Reads Them

6 minute readUpdated October 2026Explore more

TL;DR

When you send Claude a PDF, Anthropic's docs say each page is converted into an image and the page's text is extracted and sent alongside it, and you are charged for both the text tokens and the image tokens. For text-heavy documents you often only need the words. MarkItDown is a free, MIT licensed Microsoft tool with over 188,000 stars on GitHub that turns PDFs, Word, PowerPoint and Excel files into clean markdown. Convert first, hand Claude the .md file, and keep the original PDF for pages where the charts actually matter.

Learn Claude Code. Earn income. Only $9.

👉 https://www.skool.com/claudecodeclub

Why a PDF can cost you twice

This is straight from Anthropic's PDF support docs (checked October 3, 2026). When you send a PDF to Claude, "The system converts each page of the document into an image" and "The text from each page is extracted and provided alongside each page's image."

The same page explains the bill. Text: each page typically uses 1,500 to 3,000 tokens, depending on how dense the content is. Images: "Because each page is converted into an image, the same image-based cost calculations are applied." So on the API you pay for the page's words and for a picture of the page.

  • Claude API: PDFs are processed as text plus a page image, billed as text tokens plus image tokens.
  • Claude apps: Anthropic's help center says Claude analyzes both text and visual elements like images, charts and graphics in PDFs of 100 pages or fewer (from 101 to 1,000 pages it processes text only). On a subscription you are not billed per token, but larger inputs use more of your usage.
  • Claude Code: for PDFs longer than 10 pages, Claude Code reads in page ranges, and the docs say page-range reads render the pages with pdftoppm. Those pages arrive as images.
PDF support (Claude Platform docs)

Anthropic's official explanation of how PDFs are processed and priced.

What MarkItDown is

microsoft/markitdown

188,189 stars on October 3, 2026. MIT license. Current PyPI release 0.1.8.

MarkItDown is a lightweight Python utility from Microsoft for converting files to Markdown "for use with LLMs and related text analysis pipelines." It keeps the structure that matters, like headings, lists, tables and links. The README notes that markdown is close to plain text and "highly token-efficient."

  • PDF, Word, PowerPoint and Excel
  • HTML, CSV, JSON and XML
  • EPUB books and ZIP files (it goes through what is inside)
  • Images (metadata and OCR) and audio (metadata and speech transcription)
  • YouTube URLs

Install it

MarkItDown needs Python 3.10 to 3.14. Use a virtual environment, as the README recommends.

bashpython3 -m venv .venv
source .venv/bin/activate
pip install 'markitdown[all]'

Only need PDFs and Office files? pip install 'markitdown[pdf, docx, pptx, xlsx]' installs less.

Convert your files

bash# One file
markitdown report.pdf -o report.md

# A Word doc and a spreadsheet
markitdown proposal.docx -o proposal.md
markitdown pricing.xlsx -o pricing.md

# Every PDF in the current folder
for f in *.pdf; do markitdown "$f" -o "${f%.pdf}.md"; done

Then point Claude Code at the .md file instead of the PDF, for example: read report.md and summarize the key numbers.

Make it automatic in Claude Code

Option 1: add MarkItDown's official MCP server, so Claude Code can convert any file itself. It exposes one tool, convert_to_markdown, which takes a file, http or data URI.

bashpip install markitdown-mcp
claude mcp add markitdown -- markitdown-mcp

Option 2: add this rule to your project's CLAUDE.md so Claude converts before it reads.

markdown## Documents
- Before reading any PDF, Word, PowerPoint or Excel file, convert it with: markitdown "<file>" -o "<file>.md"
- Read the .md version, not the original.
- Only open the original file (and only the specific pages) when the markdown is missing something visual I need, like a chart, a diagram or a scanned page.
- Never re-convert a file that already has an up-to-date .md next to it.

Option 3: one prompt that does a whole folder.

promptInstall Microsoft's MarkItDown in a virtual environment (pip install 'markitdown[all]', see https://github.com/microsoft/markitdown). Then convert every PDF, DOCX, PPTX and XLSX file in the ./docs folder to markdown, saved next to the original with a .md extension. Report any file that came out empty or nearly empty, because those are probably scanned images. From now on in this project, read the .md versions instead of the originals.

When to keep the PDF

  • Charts, diagrams and photos. Markdown keeps the words, not the picture. If the answer is in a chart, Claude needs to see that page.
  • Scanned PDFs. A scan is a picture of text. Basic conversion finds little or nothing to extract. MarkItDown has an OCR plugin (markitdown-ocr) and Azure options, but those call a vision model or a paid Azure service.
  • Layout-heavy files. The README says the output is meant for text analysis tools and "may not be the best option for high-fidelity document conversions for human consumption."
  • Tip: convert first, then send only the specific PDF pages you still need visually. You get the words cheaply and the pictures only where they matter.

Check the savings on your own files

Token use depends on the document, so measure instead of guessing. On the API, Anthropic's token counting endpoint tells you the input tokens for a request before you send it. Count the PDF, then count the .md, and compare. Anthropic says token counting is free to use. Docs: https://platform.claude.com/docs/en/build-with-claude/token-counting

Quick start

  1. 1pip install 'markitdown[all]' in a virtual environment.
  2. 2Run markitdown yourfile.pdf -o yourfile.md.
  3. 3Give Claude the .md file. Keep the PDF for chart-heavy or scanned pages.
  4. 4Add the CLAUDE.md rule or the MCP server so it happens every time.

Now you just feed Claude the words and keep your tokens for the work. Inside the Claude Code Club we share the token-saving setups we actually run and help each other get them working. It's nine dollars a month at https://www.skool.com/claudecodeclub/about. Everything on this page works without it.

Common questions

  • Does Claude really charge for both the text and the image of each PDF page?

    On the API, yes. Anthropic's PDF docs say each page is converted into an image with its text extracted alongside, that text typically costs 1,500 to 3,000 tokens per page, and that image-based cost calculations also apply to each page.

  • Is MarkItDown free?

    Yes. It is an open-source Microsoft project under the MIT license. Conversion runs locally. Only the optional Azure Document Intelligence and Content Understanding features are billable Azure calls.

  • Will I lose charts and images?

    Yes, mostly. Markdown keeps text, headings, lists and tables. If you need Claude to read a chart or diagram, send that page as the original PDF or an image.

  • Does it work on scanned PDFs?

    Basic conversion extracts the text layer, which a scan does not have. The markitdown-ocr plugin adds OCR using a vision model you provide, and Azure Document Intelligence is a paid alternative.

  • How do I use it inside Claude Code?

    Either convert files yourself with the markitdown command, add the official markitdown-mcp server with claude mcp add markitdown -- markitdown-mcp, or add a CLAUDE.md rule telling Claude to convert documents before reading them.

Want every token-saving setup we actually run?

Get the other 4 in the cost-cutting stack, plus 8,000+ members - $9/mo, cancel anytime.

Join the Club