Extract text from scans with local OCR

Turn English text in scanned PDFs, PNGs, JPEGs, and WebP images into a searchable PDF or a TXT file. OCR runs on your device without uploading the source.

Drop your PDF or image file here or click to choose a file PDF or image file · processed on this device

Read scans without sending confidential pages to an OCR server

Use this tool when the words in a PDF are pixels rather than selectable text. It also reads English text from PNG, JPEG, and WebP images. FileGizmo loads a self-hosted Tesseract recognition engine only after you choose to process the file. The source stays in the current browser tab, and both results, a searchable PDF and a plain UTF-8 text file, are created on your device.

For a PDF, choose only the pages you need. FileGizmo renders those pages locally, sends each canvas to a browser worker, and adds the recognized words to those pages of your original file; the other pages come back unchanged. The first run downloads the English recognition model and engine from FileGizmo’s own origin. Those assets can then be cached for repeat and offline use. No cloud OCR account, API key, or remote document library is involved.

OCR is an interpretation, not a certified transcription. Clear, upright text at a useful resolution performs best. Low contrast, handwriting, skewed pages, unusual fonts, multi-column layouts, and tables can reduce accuracy. The result shows average confidence as a review signal, but confidence is not proof that a name, date, amount, or identifier is correct. In the searchable PDF a misread word is still misread, so a search can miss it.

The text file preserves page boundaries, not visual layout. Check important values against the scan before using them in a legal, medical, financial, academic, or automated workflow. If the PDF already contains selectable text, the faster PDF to Text tool can extract that text layer without running OCR.

How to OCR a scanned PDF or image locally

  1. 1

    Choose a scanned PDF or image

  2. 2

    Select PDF pages and recognition quality

  3. 3

    Download a searchable PDF or copy the extracted text

Everything you need

Scanned PDFs and common images

Render selected PDF pages locally or read a PNG, JPEG, or WebP image without creating a remote processing job.

Searchable PDF and page-by-page text

The PDF keeps every original page and adds the recognized words as an invisible text layer you can search, select, and copy. The TXT file has a heading for each page.

English model with confidence scores

This first release uses the self-hosted English model. It reports average confidence and asks you to verify names, numbers, and tables.

Nothing is uploaded, and you can verify it

Processing runs on this device. Open the browser’s Network panel while the tool works and you will see no file upload request to FileGizmo. Temporary previews use local browser URLs, and closing or resetting the page releases them instead of leaving a server copy behind.

No limits, no watermarks

Use the tool without a daily task counter, account wall, output watermark, or artificial upload cap. FileGizmo does not meter a transfer it never receives. Available memory, processor speed, browser canvas limits, and the file format itself set the practical limit.

Works on any device, and offline after first load

Use a current browser on desktop, Android, iPhone, or iPad without installing an app. Once this tool and its required assets load, the cached workflow works without a connection. Large jobs are usually more comfortable on a desktop with additional memory.

Frequently asked questions

Does FileGizmo upload my scan for OCR?

No. The OCR worker, recognition engine, and English language data load from FileGizmo and run inside your browser. The selected file is not posted to a processing endpoint.

Can this make a scanned PDF searchable?

Yes. After OCR, Download searchable PDF gives you the original pages with the recognized words added as an invisible text layer, so the scan looks the same and its words can be searched, selected, and copied. Pages that already contain selectable text keep their own text.

Which languages are supported?

This release supports English. It may read isolated Latin-script text in other languages, but FileGizmo does not claim accurate multilingual recognition yet.

Will OCR preserve tables and page layout?

The searchable PDF keeps the original pages exactly as they look, because its text layer is invisible. The TXT file is plain UTF-8 text with page headings, so columns, handwriting, faint scans, decorative fonts, and tables can need manual correction there.

Related guides

PDF/A, PDF/X, or plain PDF: which one do you need?Choose plain PDF for everyday use, PDF/A for long-term preservation, or the exact PDF/X profile your commercial printer requests, then validate the result.