PDF Nest
All guides

How to extract text from a PDF

There are two kinds of PDF, and they behave in opposite ways. One carries the words and can be copied from; the other carries a photograph of the words and cannot be copied from by anything. Telling them apart takes two seconds, and it is the first thing worth doing.

Which kind of PDF you have

Open it in any PDF reader and try to select a word with the mouse. If a word highlights, the document has text in it and its own words can be taken out — exactly, in English; very nearly, in Sinhala and Tamil. If the cursor sweeps over the page selecting nothing, or highlights the entire page at once as a single block, the page is a picture and there are no words in the file to take out.

That test is the whole answer, and it is worth doing before trying anything else, because no tool on earth extracts words from a document that has none. A scanner produces pictures of paper; so does a phone photograph saved as a PDF; so does a fax. All three look like documents on screen and none of them contains text.

Taking the words out

Drop the PDF on the tool and press the button. The text appears in a box you can select by hand, with one button to copy all of it. Nothing is installed, no account is needed, and the document is opened inside the browser tab rather than sent anywhere — which for a contract, a statement or a letter is the difference between using a tool and thinking better of it.

What comes back is plain text: the words, in reading order, with paragraph breaks where the page showed them. Headings, bold and italics are not carried across, because a PDF stores styling separately from the words and the point of this is the words.

When the PDF is a picture

Then the words have to be read rather than extracted, and that is a different job with a different answer: read the pages with Image to Text. Point it at the pages as images and it will recognise Sinhala, Tamil and English together, and tell you how sure it is of each page. The result is a draft to check, not an exact copy — which is the honest difference between recognition and extraction, and the reason the two are separate tools here.

A useful rule for a mixed document, which is common in anything old: extract first, and recognise only the pages that came back empty.

Keep the text, or keep the layout

Extraction throws the layout away, and that is usually the point — you want the words to put somewhere else. If what you actually want is the page itself, with its layout intact, that is a different tool: the page-as-PDF output of Image to Text keeps the document exactly as it looked and still lets you search it.

Do it now, free

Most PDFs are not pictures. They are words with positions attached, stored along with the mapping from each drawn shape back to the character it stands for — which is what makes text selectable in a reader at all. Reading that mapping back is not recognition: nothing is guessed at. For English it is exact. For Sinhala and Tamil it is very close and not perfect, and the section below says by how much rather than leaving you to find out.

Extract text from a PDF →

Runs in your browser. Nothing is uploaded.

Questions

Can I copy the text out by hand instead?

You can, and for one paragraph that is faster. For a page or more it is not: selecting across columns and page breaks picks up the wrong things in the wrong order, and it misses text that is in the file but not drawn where you expect it.

Why does my PDF come back as one long paragraph?

Because the page did not have paragraph gaps in a form that could be measured — usually a design with a lot of whitespace that means nothing, or a page that is mostly a table. The words are all there; the line breaks between them are the part that is a reconstruction rather than a copy.

Do I need to install anything?

No. It runs in the browser tab. You can even disconnect from the network first and it will still work, which is also the easiest way to satisfy yourself that nothing is being sent anywhere.

What about a PDF with a password?

It has to be opened before it can be read, and there is nowhere to type the password yet, so a protected document is refused rather than half-read.

All tools