Sinhala and Tamil OCR, without uploading the page
Most OCR tools can only read your page by taking it. The recognition happens on their server, so the document — a letterhead with a name, an address, sometimes a national identity number on it — is uploaded first and deleted, you are told, later. That is the part worth avoiding, and for Sinhala and Tamil it turns out to be avoidable.
Where the document goes, and how to check it yourself
Everything on this page happens inside the tab it was opened in. Your image is read into the page, compared against models that are already on your device, and never transmitted anywhere — there is no upload step, no queue and no account, because there is no server holding the document to begin with. For an English screenshot that is a convenience. For a Sri Lankan letterhead, which normally carries a name, an address and often an identity number under the crest, it is the whole point.
You do not have to take that on trust, and you should not. Two ways to settle it in under a minute: open your browser's network tab while a page is being read and watch that nothing is sent, or load the tool, then switch the network off entirely and read an image anyway. The second one is the harder test, and it passes because the recognition engine and the language models are served from this site rather than from a content network — the models arrive with the page, so recognition never needs a connection at all.
What decides whether it works: the photograph, not the language
Ask whether Sinhala or Tamil recognition is accurate and the honest answer is that it depends almost entirely on the picture, and much less on the script than you would expect. On a clean 300 dpi page in a modern typeface, read with all three models, neither script produced a single character error across fifteen samples each. Re-photograph the same page at 110 dpi through a heavy JPEG and Sinhala fell to 7% of characters correct. Same engine, same font, same words — the only thing that changed was how the page was captured.
So the advice that matters is about taking the picture. Fill the frame with the page rather than the desk around it, hold the phone square to the paper rather than at an angle, keep the light even, and keep your own shadow off the page — a shadow across a column of Sinhala is read as damage. Do not photograph a screen: a picture of a monitor adds a moiré pattern that costs more accuracy than the resolution gains. If the document is already a digital PDF, do not photograph anything — its text can be copied out exactly, with no recognition involved.
Two things are genuinely outside what this does. Handwriting, in any of the three languages: printed Sinhala and Tamil are read, written ones are not, and no setting changes that. And pages where two blocks of text sit side by side, because deciding which block is read first becomes a guess.
Getting from a photograph to a document you can send
There are two useful outputs, and they answer different problems. The first button gives you the page itself as a PDF: the photograph, with the recognised words in a layer underneath it, so the crest, the letterhead and the layout survive exactly as they were and a search still finds the words. That is the one to send when the page has to look like the page.
The second is the words themselves, in the box under the result, where you can correct anything the reader got wrong and copy it out. For a Sinhala or Tamil page that is half of a larger job, and the other half is done by the second tool. Paste the corrected text into Text to PDF and it sets real සිංහල and தமிழ் — shaping the letters with the font's own rules, rather than drawing one character after another and hoping — with margins, headings, bullets and page numbers. A photograph goes in at one end and a typed, laid-out document comes out at the other, without anyone retyping it and without the page leaving the device.
What a transcription is not
Reading a document is not the same as certifying it, and the difference matters when the result is going to an embassy, a university or a bank. Recognised text is a draft: it is very good on a clean page and it is not a certified translation, it carries no legal weight, and an authority asking for the document itself wants the page, not a transcript of it. For that, send the page-as-PDF output — it is the original page, down to the crest and the letterhead, with the words underneath only as a convenience.
The same honesty applies to the text. Where the reader is unsure it does not mark the word; it returns its best guess, which is why the confidence score is shown rather than hidden, and why a low score is named as a problem instead of quietly producing a page of convincing nonsense. Check anything you are about to rely on against the picture.
Do it now, free
A photo of a letter, a scan of a form, a picture of a receipt: what you want back is the page — the crest, the letterhead, the layout — with words inside it that can be selected and searched. That is what the first button gives you: the page as it stands, as a proper PDF, with the recognised text underneath it. Copying the words out as plain text is the other half of the job, and only worth doing when the words are all you need.
Runs in your browser. Nothing is uploaded.
Questions
Does it really read Sinhala and Tamil, or only English?
Both scripts are read with their own models, and the accuracy is measured rather than claimed: on clean 300 dpi pages in a modern typeface, neither script produced a single character error across fifteen samples each. What the models cannot do is compensate for a bad photograph — at 110 dpi through a heavy JPEG, Sinhala fell to 7% of characters correct.
Do I have to tell it which language my page is?
No. The page starts with the English model, which is the cheap reading, and if that comes back unsure the Sinhala and Tamil models are fetched and it is read again — which is what a Sri Lankan letterhead carrying all three on one sheet needs, since nobody should have to guess in advance which languages a photograph contains. A reading that did not have all three models can always be repeated with them from the button beside the result.
Is my document uploaded to a server?
No. Recognition runs in the browser tab you already have open, against models that are on your device, and there is no upload step. The way to check is to open your browser's network tab while a page is being read, or to load the tool, disconnect from the internet, and read an image anyway — it still works, because the engine and the models were served with the page.
Can I get an editable document out, or only a picture of one?
Both. One button gives you the page as a PDF with the recognised words in a selectable layer under it, which keeps the letterhead and the layout. The words are also in the box under the result, where you can correct them and paste them into Text to PDF — which sets Sinhala and Tamil properly, with margins, headings and page numbers, as a fresh typed document.
Will it read handwriting?
No, in any language. Print only — a screenshot, a scan, or a steady, well-lit, close-up photograph of printed or typed text. As a rule the camera in a phone is more than good enough; the light and the steadiness are what decide it.
Why is the first run on this page slower than the other tools?
Recognition needs a real engine inside the tab, about four megabytes, plus a language model. An English page takes the English model and stops there — about 6 MB in total, once. A page that turns out to be Sinhala or Tamil pulls in those two models as well, about 2.4 MB more, and is read again. Everything is cached afterwards, and every other tool on this site opens without waiting for any of it.
All tools
Merge PDF files
Combine several PDFs into one file, in the order you choose.
Split PDF files
Break one PDF into separate files, or pull out only the pages you need.
Delete, reorder and rotate PDF pages
Delete, reorder, rotate and duplicate pages in a visual grid.
Edit PDF files
Add text, a signature, white-out, a watermark or page numbers.
Compress PDF files
Cut file size for email and upload limits — three levels of compression.
Convert PDF to images
Export every page as a PNG or JPG, delivered as a ZIP.
Convert images to PDF
Turn JPG, PNG or WEBP images into a single PDF document.
Rotate PDF pages
Fix sideways or upside-down scans, for all pages or a few.
Extract text from an image
Read the words out of a photo or a screenshot, without retyping them.
Turn text into a PDF document
Type or paste a page and get it back as a laid-out document, not a text dump.
Extract text from a PDF
Pull the words out of a PDF as text you can copy — no upload, no account.