Skip to content
Klarfolder

Split a scanned PDF into separate documents

The feeder scanned everything at once: thirty documents in a single file. The tool proposes the cuts, shows them on a strip of thumbnails, and each one is corrected with a click.

Or drop a scanned stack in one PDF, even three hundred pages long here. Nothing is sent: everything stays in this tab.

When to use it

A stack went through the feeder and came out as one PDF: the folder exists, but no document is separated.

Or a client sent “the file”, scanned in one go, and the documents have to be pulled out one by one before anything can be done with them.

What this tool does

What it does not do: the tool does not guess what a document means. It spots breaks between two pages, and when two documents follow one another with no signal at all it misses one: that is measured, and it is fixed with a click.

What is read in the document

No file is uploaded: the folder is read by your browser, in this tab, and the page is not allowed to open a connection to any other site.

Frequently asked questions

How many cuts does the tool miss?
Measured over ten synthetic stacks, 1,048 pages and 639 documents: 98.3 % of the cuts are found, and none is spurious. Eleven corrections in all, four at worst on one stack.
What happens when two documents follow one another with no signal?
That is the limiting case, and it is in the measurement: a whole stack of the corpus is made of documents with no repeated header, no footer and no numbering. The tool finds 96.6 % of the cuts there, and the two it misses take a click each.
Three hundred pages, does that hold?
Yes. Reading the text of a 300-page stack takes under a second; reading a fully scanned stack takes 35 seconds with six workers.
Do the documents keep their quality?
Yes: pages are copied, not redrawn. A signed page stays what it was, and a text layer stays readable.

The other tools