Extract Text from a Scanned PDF

Read printed text from PDF pages in your browser, then copy or download the recognised text as a plain .txt file. This tool does not create a searchable PDF.

Extract Text with OCR

Upload a PDF and extract plain text with OCR

How to OCR a scanned PDF — 3 steps

  1. 1

    Add your scanned PDF

    Open the OCR editor and choose one PDF. It works page by page on scans and other PDF pages that do not already contain useful text.

  2. 2

    Run text recognition

    Choose a language and start recognition. The browser renders each page as an image and Tesseract reads the visible characters.

  3. 3

    Copy or download the text

    Review the plain-text result, copy it to the clipboard, or download it as a .txt file. This tool does not create a searchable PDF.

What this PDF OCR tool actually produces

The editor renders each PDF page, sends the page image to Tesseract OCR, and collects the recognized characters as plain text. It adds page separators such as "--- Page 1 ---" so the extracted text can be matched back to the source.

The result appears in an editable text area. You can copy it or download a text file, but the original PDF is not modified and no searchable PDF or invisible text layer is created. Tables, columns, spacing, images, and page layout may be simplified or lost.

Recognition depends on the source page and the selected language. Clear, upright printed text generally works best; blur, skew, shadows, handwriting, unusual fonts, and mixed languages can produce mistakes. Always check names, numbers, and important clauses against the scan.

Languages, limits, and what to check

The editor offers English, Spanish, French, German, Simplified Chinese, Japanese, Portuguese, and Russian. Select the language that matches most of the page before starting OCR.

OCR processes every page in the PDF and can take longer as page count and image size increase. It is intended for text extraction, not handwritten-text transcription, layout-preserving conversion, or a guaranteed word-for-word result.

Compare the text with the original before using it in a contract, application, record, or other important document. Keep the source PDF because the downloaded text file is only the recognition output.

Your files never leave your device

The PDF is rendered and read in your browser. The file and the recognized text are not uploaded to our server.

Most online PDF sites upload your document to their servers, process it there, and keep a copy for minutes, hours, or longer. PilotPDF works differently: the page downloads a small processing engine into your browser, and your file is opened and edited in your device's memory using WebAssembly. Nothing is transmitted, so there is nothing for a server to store, leak, or scan. You can even disconnect from the internet after the page loads and the tool keeps working.

Read more about how this works in our privacy-by-design explainer.

Related tools

Frequently Asked Questions

No. This editor extracts recognised characters into an editable text area and downloads a plain .txt file. The uploaded PDF is not rewritten.
The result downloads as extracted-pdf-text.txt. It includes page separators so you can tell where each page's text came from.
English, Spanish, French, German, Simplified Chinese, Japanese, Portuguese, and Russian are available in the language menu.
No. The output is plain text, so columns, tables, spacing, images, and page design may not be preserved.
It is intended mainly for printed text. Handwriting, unusual fonts, blur, shadows, and skewed pages can lead to poor results.
There is no fixed accuracy promise. Clear, upright pages usually produce better text than low-quality scans, so check important names, numbers, and wording against the PDF.
Yes. The result is shown in an editable text area before you copy it or download the text file.
Yes. The editor runs recognition across all pages in the selected PDF.
Yes. The PDF is rendered and recognised in your browser without being uploaded.