FileGizmo — PDF tools

Extract text from scans with local OCR

Turn English text in scanned PDFs, PNGs, JPEGs, and WebP images into a searchable TXT file. OCR runs on your device without uploading the source.

Never uploadedYour file stays on this device
Drop your PDF or image here or click to choose a file PDF or image file · processed on this device

Why this tool

Read scans without sending confidential pages to an OCR server

Use this tool when the words in a PDF are pixels rather than selectable text. It also reads English text from PNG, JPEG, and WebP images. FileGizmo loads a self-hosted Tesseract recognition engine only after you choose to process the file. The source stays in the current browser tab, and the result is created as a plain UTF-8 text file on your device.

For a PDF, choose only the pages you need. FileGizmo renders those pages locally and sends each canvas to a browser worker. The first run downloads the English recognition model and engine from FileGizmo’s own origin. Those assets can then be cached for repeat and offline use. No cloud OCR account, API key, or remote document library is involved.

OCR is an interpretation, not a certified transcription. Clear, upright text at a useful resolution performs best. Low contrast, handwriting, skewed pages, unusual fonts, multi-column layouts, and tables can reduce accuracy. The result shows average confidence as a review signal, but confidence is not proof that a name, date, amount, or identifier is correct.

The output preserves page boundaries, not visual layout. Check important values against the scan before using them in a legal, medical, financial, academic, or automated workflow. If the PDF already contains selectable text, the faster PDF to Text tool can extract that text layer without running OCR.

Simple by design

How to OCR a scanned PDF or image locally

  1. 1

    Choose a scanned PDF or image

  2. 2

    Select PDF pages and recognition quality

  3. 3

    Review copy or download the extracted UTF-8 text

Built for the whole job

Everything you need

Scanned PDFs and common images

Render selected PDF pages locally or read a PNG, JPEG, or WebP image without creating a remote processing job.

Page-by-page output

PDF results include clear page headings so you can trace extracted text back to the source scan.

Honest English recognition

This first release uses the self-hosted English model. It reports average confidence and asks you to verify names, numbers, and tables.

Nothing is uploaded — verify it yourself

Processing runs on this device. Open the browser’s Network panel while the tool works and you will see no file upload request to FileGizmo. Temporary previews use local browser URLs, and closing or resetting the page releases them instead of leaving a server copy behind.

No limits, no watermarks

Use the tool without a daily task counter, account wall, output watermark, or artificial upload cap. FileGizmo does not meter a transfer it never receives. Available memory, processor speed, browser canvas limits, and the file format itself set the honest practical boundary.

Works on any device — and offline after first load

Use a current browser on desktop, Android, iPhone, or iPad without installing an app. Once this tool and its required assets load, the cached workflow works without a connection. Large jobs are usually more comfortable on a desktop with additional memory.

Good to know

Frequently asked questions

Does FileGizmo upload my scan for OCR?

No. The OCR worker, recognition engine, and English language data load from FileGizmo and run inside your browser. The selected file is not posted to a processing endpoint.

Can this OCR a scanned PDF?

Yes. FileGizmo renders the PDF pages you select into local canvases, then recognizes the pixels on your device.

Which languages are supported?

This release supports English. It may read isolated Latin-script text in other languages, but FileGizmo does not claim accurate multilingual recognition yet.

Will OCR preserve tables and page layout?

No. The output is plain UTF-8 text with page headings. Columns, handwriting, faint scans, decorative fonts, and tables can require manual correction.

Learn more

Related guides

PDF guidePDF/A, PDF/X, or plain PDF: which one do you need?Choose plain PDF for everyday use, PDF/A for long-term preservation, or the exact PDF/X profile your commercial printer requests, then validate the result.