โ† All guides ยท July 12, 2026

How to Extract Text From a PDF (Even Scanned Ones)

PDFs are everywhere, and getting text out of them is sometimes trivial and sometimes surprisingly stubborn. The trick is knowing which kind of PDF you have, because that decides the method.

The two kinds of PDF

Digital (text-based) PDFs are created from a document, like exporting from Word. The text is already real, selectable characters. You can usually highlight a sentence with your cursor.

Scanned (image-based) PDFs are pictures of pages, often from a scanner or a photo. There is no real text inside, just an image. If you try to select text and nothing highlights, you have a scanned PDF, and you will need OCR to read it.

Knowing which you have saves you a lot of confusion.

Extracting from a digital PDF

If you can select the text, extraction is easy:

  1. Open the PDF and select the text you want, or all of it.
  2. Copy and paste it where you need it.

For a whole document or to get a clean export, run it through our PDF to text converter, which pulls the text out for you and hands back a tidy file. No sign-up, and the file is deleted after processing.

Extracting from a scanned PDF

Selecting does nothing on a scanned PDF, so OCR has to read the page images. Use the scanned PDF to text tool, which runs recognition across each page and returns editable text.

Because this relies on OCR, the same quality rules apply as with any image: a crisp, high-contrast, straight scan reads far better than a faint or skewed one. The principles in our guide on how to improve OCR accuracy apply directly here.

Step by step

  1. Check the PDF type by trying to select text.
  2. Upload it to the PDF to text tool (or scanned PDF to text for image-based files).
  3. Wait while it processes; multi-page files take a little longer.
  4. Review the output, especially numbers and names, since OCR is never flawless.
  5. Copy or download the text, or export to an editable Word document if you need to keep editing.

Common questions

How do I know if my PDF is scanned?

Try to select a line of text with your cursor. If it highlights, it is a digital PDF with real text. If nothing selects and the page behaves like an image, it is scanned and needs OCR.

Will the formatting survive?

Simple single-column documents come through well. Multi-column layouts and tables are the hardest for any OCR engine, so check those areas and expect to tidy them up. Reading order can occasionally shuffle on complex pages.

Can I extract just a few pages?

You can crop or split the PDF to the pages you care about before uploading, which also speeds up processing and reduces clutter in the result.

Get your text

Have a PDF to crack open? Send a digital one to PDF to text, or a scanned one to scanned PDF to text, and you will have editable text in moments. For images rather than PDFs, our extract text from an image guide covers that path.

Try it now

Extract text from any image or PDF โ€” free, no sign-up.

Open the converter