PDF to Markdown Converter

Convert a PDF to Markdown in this browser tab, with a local AI model for scanned pages. Nothing uploaded.

ZERO UPLOAD · ALL LOCAL
  1. Drop one PDF onto the drop zone, or click it to choose a file.
  2. Keep All pages to convert everything, or pick Selected pages and type pages and ranges such as 1-5, 8.
  3. If the PDF has scanned pages, read the info line: it says whether they will use the AI model (and whether it needs downloading) or basic OCR.
  4. Click Convert to Markdown and watch the Model and Converting progress rows.
  5. Step through the Flagged pages, compare each source page with its Markdown, and use Re-run with AI model on a flagged text page if you want a second pass.
  6. Switch between Markdown and Preview, tick Include blank pages or Include page markers if you need them, then Download .md or Copy the result.

PRIVACY GUARANTEED Your PDF never leaves this tab. Text extraction, the AI model, and OCR all run on your device.

No file selected

Pages

Two kinds of PDF page, two ways to read them

A PDF page either carries its own text or it carries a picture of text, and the difference decides everything about conversion. Word processors, browsers, and report generators write a text layer: every word sits in the file with its font, size, and position, so a converter can read it back exactly. Scanners and phone scan apps write an image, and the words exist only as pixels. Copy and paste works on the first kind and returns nothing on the second.

Inside one document you often get both. A contract might be typed on pages one to nine and scanned on page ten where somebody signed it. When you drop the file, this tool checks every page for a usable text layer, meaning at least 20 characters of real text, and the file row shows the result, for example 14 pages · 12 with a text layer. Pages with text go through the fast path, pages without it go to the scanned path, and you never pick a mode by hand.

Rebuilding structure from the text layer

A text layer stores words and coordinates, not headings or lists. Markdown needs structure, so the converter infers it. Through pdf.js, the same PDF engine Firefox uses to display PDFs, the tool reads each text item with its position and font size, rebuilds lines from items that share a baseline, and measures the most common size on the page as body text.1 Lines clearly larger than body text become headings, with the largest size mapped to #, the next to ##, and anything smaller to ###.

Paragraphs come from spacing. When the gap between two lines is well above the page's typical line gap, a new paragraph starts; otherwise the lines join, and a word split with a hyphen at a line end is put back together. Lines that open with a bullet glyph or a number such as 1. become list items, and extra indentation nests them. Runs in a monospace font turn into a fenced code block that keeps its relative indentation. None of this needs a model. It runs in milliseconds per page, and the output uses the exact words stored in your PDF.

When the words form a grid

Tables have no special marker in a text layer either. On each line, the converter treats a wide horizontal gap between words as a cell boundary, and when at least three rows break at the same three or more positions, it reads those positions as table columns. Every line that fits the grid becomes a row of a Markdown pipe table, with the first row as the header, so a price list or a timetable keeps its columns instead of collapsing into one long sentence. Because this is still an inference from positions alone, the page also gets a Table detected flag, and it pays to compare the table against the source page before you rely on it.

Scanned pages and the local vision model

Pixels need a reader. For scanned pages the tool runs granite-docling-258M, a 258-million-parameter vision-language model that IBM released under the Apache 2.0 license, which looks at a page image and writes DocTags, a compact markup that labels each title, paragraph, list, code block, formula, and table together with its position on the page.23 The tool then turns those DocTags into Markdown: tables become pipe tables, merged cells stay as HTML, and running page headers and footers drop out. Before the model sees a page, the tool renders it with its longest side at 2048 pixels, and the model's processor cuts that image into 512-pixel tiles plus one reduced overview, so small print stays legible.34

Precision matters more than download size here. Through Transformers.js on your graphics card via WebGPU, the tool loads the model's full 32-bit build of about 1.27 GB, because the compressed builds broke down in testing on a consumer GPU: the 16-bit files produced nothing but exclamation marks, and the 4-bit files dropped a whole table column while the rest of the page still looked plausible.56 After the first download, your browser keeps the files in its cache, so a later visit loads the model in seconds.

Speed and the basic OCR fallback

Speed depends on your hardware and on how much text a page holds. A recent discrete GPU handles a page far faster than an integrated one, and a dense page of small print takes longer than a short letter, so the Converting row shows an estimated time left once the first scanned page finishes. When your browser has no WebGPU, the download fails, or the model returns unreadable output, scanned pages go to Tesseract instead.7 That fallback keeps Tesseract's own paragraph and line structure, joins wrapped lines into sentences, and separates paragraphs with a blank line, but it cannot rebuild tables or headings, so the info line says which path your pages will take before you click.

Pages worth a second look

Every converter guesses somewhere. Instead of hiding those guesses, this tool marks the pages where it had a concrete reason to doubt itself and writes the reason next to the source page. Text pages get flagged when the words line up in a grid, which usually means a table; when the lines split into two columns; or when a large share of characters are unreadable, a sign that the PDF's fonts map glyphs to the wrong characters.8 Two-column pages are still read column by column, and the flag simply asks you to confirm the order.

Model and OCR pages have their own warnings. A model page gets flagged when the output reached the length limit and may be cut off, or when the text started repeating, a known failure mode of generative models.9 An OCR page gets flagged when Tesseract's average confidence falls below 70 percent.10 With the Flagged filter on, the arrows step only through those pages, and on a flagged text page with WebGPU available you can press Re-run with AI model to replace that one page with the model's reading.

What stays on your device

Your PDF is opened from memory in this tab and never sent anywhere. That matters because the documents people convert to Markdown are often exactly the ones they should not paste into an online service: contracts, medical letters, internal reports, or notes headed for a private wiki. The model files are the only download, they come from Hugging Face or from CapyToolkit's own CDN if Hugging Face is slow to answer, and your browser keeps them in its cache for later visits.

Shaping the final Markdown file

Tables with merged cells stay as HTML inside the Markdown, because pipe tables cannot express a cell that spans two rows, and CommonMark allows raw HTML blocks.1112 Photos and charts are not exported as images: on a model page, a figure becomes a hidden HTML comment with the word image, followed by its caption as plain text, and a text page simply skips pictures. With Include page markers on, each page starts with an HTML comment naming its page number, which Markdown renderers hide but you can search for when checking a passage against the original.

Empty pages get their own switch. When the model or OCR finds nothing on a scanned page, the tool labels it Blank page and leaves it out of the file, which suits the separator sheets and blank backs that document scanners pick up. Turn on Include blank pages and each of those pages keeps a page marker, so the numbering in your Markdown still lines up with the PDF. Both options update the result instantly without converting again. After that, Copy puts the whole document on your clipboard, and Download .md saves it with a name that starts with capytoolkit-, ready for Obsidian, a documentation repo, or a prompt you assemble yourself.

Sources
  1. 1.

    Mozilla, "PDF.js: PDF Reader in JavaScript," github.com, accessed September 2026. https://github.com/mozilla/pdf.js

  2. 2.

    IBM, "Granite Docling," ibm.com, accessed September 2026. https://www.ibm.com/granite/docs/models/docling

  3. 3.

    Ahmed Nassar et al., "SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion," arXiv, 2025, pp. 1–24. https://arxiv.org/html/2503.11576v1

  4. 4.

    Hugging Face, "Idefics3," huggingface.co, accessed September 2026. https://huggingface.co/docs/transformers/en/model_doc/idefics3

  5. 5.

    Hugging Face, "Transformers.js," huggingface.co, accessed September 2026. https://huggingface.co/docs/transformers.js/index

  6. 6.

    W3C GPU for the Web Working Group, "WebGPU," w3.org, September 2026. https://www.w3.org/TR/webgpu/

  7. 7.

    Tesseract OCR, "Tesseract User Manual," tesseract-ocr.github.io, accessed September 2026. https://tesseract-ocr.github.io/tessdoc/

  8. 8.

    Docling Project, "PDF text extraction fails for subsetted/custom fonts," github.com, September 2025. https://github.com/docling-project/docling/issues/2170

  9. 9.

    Ari Holtzman et al., "The Curious Case of Neural Text Degeneration," arXiv, 2019. https://arxiv.org/abs/1904.09751

  10. 10.

    Tesseract OCR, "tesseract: Advanced API," tesseract-ocr.github.io, accessed September 2026. https://tesseract-ocr.github.io/tessapi/4.0.0/a01625.html

  11. 11.

    GitHub, "GitHub Flavored Markdown Spec," github.github.com, 2019. https://github.github.com/gfm/

  12. 12.

    John MacFarlane, "CommonMark Spec," spec.commonmark.org, 2024. https://spec.commonmark.org/0.31.2/

FAQ