Extract Text from a Scanned PDF
Read printed text from PDF pages in your browser, then copy or download the recognised text as a plain .txt file. This tool does not create a searchable PDF.
Extract Text with OCR
Upload a PDF and extract plain text with OCR
How to OCR a scanned PDF — 3 steps
- 1
Add your scanned PDF
Open the OCR editor and choose one PDF. It works page by page on scans and other PDF pages that do not already contain useful text.
- 2
Run text recognition
Choose a language and start recognition. The browser renders each page as an image and Tesseract reads the visible characters.
- 3
Copy or download the text
Review the plain-text result, copy it to the clipboard, or download it as a .txt file. This tool does not create a searchable PDF.
What this PDF OCR tool actually produces
The editor renders each PDF page, sends the page image to Tesseract OCR, and collects the recognized characters as plain text. It adds page separators such as "--- Page 1 ---" so the extracted text can be matched back to the source.
The result appears in an editable text area. You can copy it or download a text file, but the original PDF is not modified and no searchable PDF or invisible text layer is created. Tables, columns, spacing, images, and page layout may be simplified or lost.
Recognition depends on the source page and the selected language. Clear, upright printed text generally works best; blur, skew, shadows, handwriting, unusual fonts, and mixed languages can produce mistakes. Always check names, numbers, and important clauses against the scan.
Languages, limits, and what to check
The editor offers English, Spanish, French, German, Simplified Chinese, Japanese, Portuguese, and Russian. Select the language that matches most of the page before starting OCR.
OCR processes every page in the PDF and can take longer as page count and image size increase. It is intended for text extraction, not handwritten-text transcription, layout-preserving conversion, or a guaranteed word-for-word result.
Compare the text with the original before using it in a contract, application, record, or other important document. Keep the source PDF because the downloaded text file is only the recognition output.
Your files never leave your device
The PDF is rendered and read in your browser. The file and the recognized text are not uploaded to our server.
Most online PDF sites upload your document to their servers, process it there, and keep a copy for minutes, hours, or longer. PilotPDF works differently: the page downloads a small processing engine into your browser, and your file is opened and edited in your device's memory using WebAssembly. Nothing is transmitted, so there is nothing for a server to store, leak, or scan. You can even disconnect from the internet after the page loads and the tool keeps working.
Read more about how this works in our privacy-by-design explainer.