PDF to Markdown
Extract PDF text as Markdown with best-effort structure. Complex layouts, tables, lists, and code formatting may need cleanup.
npx mdedit-cli convert report.pdf --to mdSame engine as this page. The command signs you in on first run.
Related conversions
Switch formats or reverse the workflow with a related converter.
How to use PDF to Markdown
- 1Upload a .pdf file.
- 2Select “Convert” to build the .md file.
- 3Copy the .md result, or download it as a file.
Details
What PDF to Markdown can extract
This converter is for text-based PDFs. It extracts readable text and returns an editable Markdown document, but PDF is a fixed-layout format rather than a structured Markdown source.
- Paragraph text usually remains readable.
- A visual heading can become plain text instead of a Markdown
#heading. - Bullet markers, list nesting, and reading order can change.
- Multi-column pages, headers, footers, and positioned text may be flattened.
Scanned PDFs and OCR
The converter does not run OCR. A scanned PDF needs a searchable text layer before it can be converted here.
If selecting and copying text in a PDF viewer produces nothing, run OCR in a dedicated scanning tool first. An image-only scan may return no usable text or a conversion error.
Tables, images, and code
- Tables: cell text may survive, but rows and columns are not reliably rebuilt as Markdown tables.
- Images: embedded figures are not exported as a bundled Markdown image folder.
- Code: characters may survive, but inline backticks, fenced code blocks, indentation, and syntax labels are not reliably inferred.
Plan to compare the result with the PDF and repair important structure after conversion.
Realistic PDF conversion example
Example: a quarterly report contains a visual title, a two-column metrics table, and a boxed SQL sample. The converted Markdown may keep the title, metric labels, values, and SQL text while flattening the table and omitting Markdown heading and code-fence markers.
That result is useful for recovering editable text, but it is not a layout-faithful PDF reconstruction.
Privacy and known limits
PDF files are uploaded to mdedit.ai servers for conversion. Uploaded files and results are automatically deleted after 30 days.
The file limit is 10 MB. Password-protected, malformed, image-only, or unsupported PDFs may fail. Review sensitive documents before uploading and verify the Markdown before publishing it.
Example
Convert PDF to Markdown. This is regular document content—no Markdown needed inside the file.
- Milestones: Draft, Review
- Next: Publish to web
| Item | Owner |
|---|---|
| Draft | Ada |
| Review | Grace |
# Release Notes The importer now streams large files instead of buffering them in memory. ```bash mdedit convert notes.md --to pdf ``` | Change | Status | | --- | --- | | Streaming import | Shipped | | Retry on 429 | In review |
Last reviewed
FAQ
- What carries over from .pdf to .md?
- Structure maps across: headings, paragraphs, lists, links, tables, code blocks, and inline emphasis. Anything that depends on the source presentation layer can be simplified or dropped, including page layout, fonts and colors, comments and tracked changes, and deeply nested or merged tables. Review those areas before you publish the result.
- What are the file limits?
- Accepted file types are .pdf. Each file must be 10 MB or smaller.
- Is my content private?
- Processed on mdedit.ai servers. Uploaded files are stored in a private S3 bucket (not publicly readable). Uploads use signed URLs, and results are provided via time-limited signed download links. Files are automatically deleted after 30 days.