Extract text from a PDF
Runs on your device — nothing is uploadedMost PDFs are not pictures. They are words with positions attached, stored along with the mapping from each drawn shape back to the character it stands for — which is what makes text selectable in a reader at all. Reading that mapping back is not recognition: nothing is guessed at. For English it is exact. For Sinhala and Tamil it is very close and not perfect, and the section below says by how much rather than leaving you to find out.
This reads a PDF that has text in it. It does not read a scan.
ReadsA PDF exported from a word processor, a browser, or this site. The words are characters in the file and come back as characters — exactly, in English; very nearly, in Sinhala and Tamil.
Does not readA scan or a photograph saved as a PDF. Inside it is a picture of a page, not words. Nothing can be extracted from words that were never there.
How it works
- Add the PDFDrop the document. It is opened in the tab — nothing is sent anywhere, and one is enough.
- Press Extract textThe words come out as they were typed, in whatever script the document uses, with no recognition and nothing to check.
- Copy the textIt lands in a box you can select by hand, with one button to copy the lot.
Why this is exact where recognition is not
There is nothing to guess at. A PDF that came out of a word processor carries the characters themselves — not shapes, not a picture, but the actual letters — together with a table that says which drawn mark stands for which character. That table exists so a reader can select and copy text, and it is the same table this tool reads.
Which is why there is no confidence percentage on this page. Recognition has to decide what a shape probably is; this does not have to decide anything. The words in the box are the words in the file.
What comes back, and in what order
A PDF does not store lines. It stores fragments of text at coordinates, and the lines you see on screen are an arrangement the reader works out. So the fragments are grouped by height to make lines, ordered across to make reading order, and a space is put back wherever the gap between two fragments is wide enough to have been a word space. Paragraph breaks come from gaps clearly wider than the line spacing.
The words are exact; the arrangement is a reconstruction. For an ordinary single-column document it matches the page closely. Where the page has several columns of text side by side, the lines share heights across the columns and come out interleaved — the honest answer is that a two-column layout does not survive the trip and should be read as a single flow.
Sinhala and Tamil, and how close this really gets
These two scripts are where extraction stops being a copy and becomes very nearly one. The characters do come out of the document's own text layer, and a Sinhala page read this way is much closer to the original than the same page recognised from a photograph — but it is not identical, and the difference is worth stating rather than leaving you to find it.
A PDF does not store text. It stores drawn shapes, plus a table saying which character each shape stands for. In Sinhala and Tamil a single character is drawn from several shapes, and a vowel sign is drawn before the letter it belongs to. Where one shape stands for more than one character the table can only name one of them, so reading the page back can reorder a vowel sign or return a neighbouring letter. Measured on a one-page Sinhala document produced by this site's own Text to PDF tool: two characters wrong out of 225, and of the sort that make a word look odd rather than unreadable.
For a scan — a photograph of a Sinhala page, or a PDF made from one — there is nothing to extract at all, and this tool says so instead of handing back an empty page. Image to Text is the tool for those, and it is worth knowing which of the two you have before you start: if you can select a word in your PDF reader, you have text; if you cannot, you have a picture.
The document never leaves the tab
Opening a PDF, reading its text and showing it to you happens entirely in the browser. There is no upload and no account, which matters more here than anywhere else on this site: the documents people need the text out of are contracts, invoices, statements and letters, and the whole reason for wanting the words is usually to work with them somewhere else.
Questions
Do two-column pages come out in the right order?
No, and it is worth knowing before you rely on it. On a page with two columns of text side by side, the lines share heights across both columns, so the text comes out interleaved. Single-column documents — which most letters, reports and invoices are — come out in order.
My PDF came back empty. Is the tool broken?
No — the document has no text in it. If you cannot select a word in your PDF reader, nobody can extract words from it, because there are none: there is a picture of a page. Read that with Image to Text instead.
Is the text exact?
In English, exactly — there is no recognition here, so nothing is guessed at and there is no confidence percentage. Sinhala and Tamil are the honest exception. A PDF records which drawn shape stands for which character, and these scripts draw one character out of several shapes, so a vowel sign can come back in a different order and now and then a letter comes back as the wrong one. On a one-page Sinhala test document two characters out of 225 were wrong. That is far better than recognising a photograph of the same page, and it is not the same as exact.
Does it keep bold, italic or headings?
Just the words, in reading order. Styling is a separate thing in a PDF and is not carried across, so the result is plain text.
Does it work on a PDF I protect with a password?
Not without the password. Opening it needs the password, and the tool has nowhere to ask for one yet.
Is my document uploaded?
No. The file is opened in the browser tab and never leaves it. There is no account, no queue, and nothing stored afterwards.
Related guides
Other tools
Merge PDF files
Combine several PDFs into one file, in the order you choose.
Split PDF files
Break one PDF into separate files, or pull out only the pages you need.
Delete, reorder and rotate PDF pages
Delete, reorder, rotate and duplicate pages in a visual grid.
Edit PDF files
Add text, a signature, white-out, a watermark or page numbers.
Compress PDF files
Cut file size for email and upload limits — three levels of compression.
Convert PDF to images
Export every page as a PNG or JPG, delivered as a ZIP.
Convert images to PDF
Turn JPG, PNG or WEBP images into a single PDF document.
Rotate PDF pages
Fix sideways or upside-down scans, for all pages or a few.
Extract text from an image
Read the words out of a photo or a screenshot, without retyping them.
Turn text into a PDF document
Type or paste a page and get it back as a laid-out document, not a text dump.