PDF to Text

Pull selectable text out of a PDF. Copy it or save as .txt - all in your browser.

Drop a PDF here

or click to browse

Select file
PDF
Your file data uploaded

0 bytes

Processed on this device
0 bytes
Files opened here
0
Requests carrying a file
0

Something not working as it should? Report an issue with PDF to Text

About PDF to Text

Pulling the text out of a PDF is the quickest way to get quotable content, feed a document into another tool, or check what is actually written in a file rather than what appears on screen.

The extraction reads the text objects stored in the document and returns them in reading order. Copy the result or save it as a plain .txt file. Because it reads what is genuinely in the file, extraction is exact rather than a best guess.

There is an important distinction to understand first. A PDF made from a word processor contains real text and extracts perfectly. A PDF made by scanning paper contains only page images, and there is no text to extract no matter which tool you use. If nothing comes out, that is what has happened, and OCR PDF is the answer because it recognises characters in the images and adds a text layer.

Multi-column layouts and complex tables are where extraction gets approximate. The text is stored positionally rather than as flowing paragraphs, so a two column academic paper can interleave columns. PDF to JSON returns each text run with its coordinates if you need to reconstruct layout properly.

How to pdf to text

  1. Open the PDF

    Drop the file in. It is parsed in your browser.

  2. Review the text

    The extracted text appears on screen so you can check it came out sensibly.

  3. Copy or save

    Copy it to the clipboard or download it as a .txt file.

Frequently asked questions

Nothing was extracted. Why?

Almost certainly because the PDF is a scan. It contains page images rather than text objects, so there is nothing to read. Run OCR PDF to recognise the characters and add a searchable text layer, then extract.

Why is the text out of order?

Text is stored by position on the page, not as flowing paragraphs. Multi-column layouts can interleave when read linearly. PDF to JSON gives you coordinates so layout can be reconstructed.

Does this preserve formatting?

No. The output is plain text, so bold, italics, headings and tables are not represented. That is usually what you want for feeding into another tool.

Can I extract from a password-protected PDF?

Remove the password first with Protect PDF, which also runs locally, then extract from the unlocked file.