FileGizmo PDF tools

Split a PDF where a phrase appears

Break a stack of documents apart wherever a phrase appears, such as an invoice number or a statement heading. Nothing is uploaded.

Never uploadedYour file stays on this device
Drop your PDF here or click to choose a file PDF file · processed on this device

Simple by design

Split PDF by Text in three steps

A great many PDFs are really several documents in a stack. A month of invoices exported as one file. A batch of statements from a portal that only offers a single download. A set of certificates, payslips or delivery notes printed to one PDF because that was the only option offered.

Every one of those has a phrase that starts each document, such as an invoice number, a statement heading or a certificate title, and that phrase is a better boundary than any page count, because the documents are not all the same length.

How the match works

Every page is searched for the phrase, and each page that contains it begins a new document. Case is ignored. The text of a page is joined together before searching, because a PDF stores text in runs, and a phrase of two words is often two runs that would never match if each were searched on its own.

Pages before the first match become a part of their own rather than being attached to the first document. That is usually a cover sheet or a summary, and it belongs on its own rather than at the front of somebody else’s invoice.

What it needs from the file

The document has to have real text. A scan is a picture of words, and no amount of searching will find a phrase in it. If searching for the phrase in your own PDF reader finds nothing, this will find nothing either, and running the file through OCR first gives it a text layer and makes it searchable.

When no page matches, the tool says so and names the phrase it looked for, rather than handing back one file and calling it a split. A phrase that appears on every page would produce one file per page, which is worth thinking about before choosing something as common as the company name.

Nothing is re-rendered

Each part is built by copying the original pages into a new document. Text stays selectable, images are not re-encoded, and the parts together are about the size of the original.

Nothing leaves the device

The pages are read and the parts are built by a worker inside your browser. There is no upload, no account, and no server copy, which is the point when the stack you are splitting is a year of somebody’s statements.

  1. 1

    Choose a searchable PDF

  2. 2

    Enter the words that start each document

  3. 3

    Split and download the parts

Good to know

Frequently asked questions

What is this for?

A single PDF holding many documents, such as a run of invoices, a batch of statements or a set of certificates, where each one begins with the same wording.

Is the match case sensitive?

No. Case is ignored, so Invoice and INVOICE both match.

Does it work on scans?

Only if the scan has a text layer. Run it through OCR first if searching in your own reader finds nothing.

Does the file leave my computer?

No. The pages are read by a browser worker on your device, with no upload, account, or server copy.

Learn more

Related guides

PDF guideHow to split a huge PDF into chaptersPlan chapter boundaries, split a large PDF into clearly named page-range files, and verify bookmarks, links, signatures, and page numbers before sharing.