Recognise text in scanned PDFs and images in Arabic, French, English and more.
Run OCR on scans and photos in three simple steps.
Drop a scanned PDF, or a JPG, PNG or WebP photo of a document.
Pick the language of the text, or a combination such as Arabic + English. For PDFs you can also limit the pages.
Wait while the text is recognised, then copy it or save it as a .txt file.
Recognition runs in your browser. Your file is never uploaded to a server.
Over a dozen languages, including mixed-language options for bilingual documents.
No account, no watermark and no daily quota. You can choose just the pages you need.
Handles multi-page scanned PDFs as well as single photos of papers, receipts and notes.
It uses Tesseract, a well-known open-source OCR engine compiled to run inside your browser. Each PDF page is turned into an image and the text in it is recognised on your own device.
No. Your PDF or image is processed locally and never leaves your device. The OCR engine and the data for the language you choose are downloaded once from a public CDN and cached by your browser, but your own document is not sent anywhere.
English, Arabic, French, Spanish, German, Italian, Portuguese, Dutch, Turkish, Russian, Hindi, Chinese (Simplified), Japanese and Korean, plus combinations such as Arabic + English, French + English and Arabic + French for mixed-language documents.
Use a sharp, straight, well-lit scan of at least 200 to 300 DPI, choose the correct language, and avoid heavy shadows. Printed text works much better than handwriting, which this tool does not recognise reliably.
The first time you use a language, its data (a few megabytes) has to download. After that it is cached, and recognition takes a few seconds per page depending on your device. For long PDFs you can enter a page range to process only the pages you need.
No. If your PDF already contains real, selectable text you do not need OCR. Use our PDF to Text tool instead — it is instant and does not require any model download.