Split a scanned PDF into separate documents
The feeder scanned everything at once: thirty documents in a single file. The tool proposes the cuts, shows them on a strip of thumbnails, and each one is corrected with a click.
Or drop a scanned stack in one PDF, even three hundred pages long here. Nothing is sent: everything stays in this tab.
When to use it
A stack went through the feeder and came out as one PDF: the folder exists, but no document is separated.
Or a client sent “the file”, scanned in one go, and the documents have to be pulled out one by one before anything can be done with them.
What this tool does
- No separator sheet is needed. The tool looks for what CHANGES between two pages: internal numbering restarting at 1, a header with not one word in common, the document type announced at the top.
- And for what proves a document continues: a sentence cut in two, a footer that repeats, a section heading (“PAR CES MOTIFS”, “Article 3”). No issuer ever opens a document with a closing formula.
- Every cut says why, in one sentence, under the thumbnail. A bare score cannot be argued with.
- One click on the line between two pages adds or removes a cut. That is the only gesture on the screen.
- Blank pages and separator sheets are set aside, shown struck through, never deleted silently.
- The documents that come out go through everything else: type identification, naming, filing, list of documents.
What it does not do: the tool does not guess what a document means. It spots breaks between two pages, and when two documents follow one another with no signal at all it misses one: that is measured, and it is fixed with a click.
What is read in the document
- The text of each page, and above all its header and footer: that is where the issuer's name and the pagination sit.
- Internal numbering (“page 2 of 4”), which is the surest signal: the only one the issuer wrote on purpose to say where their document ends.
- Pages without text are read in this tab, if you leave the box ticked.
- Pages are copied as they are into the documents: text stays text, images keep their resolution.
No file is uploaded: the folder is read by your browser, in this tab, and the page is not allowed to open a connection to any other site.
Frequently asked questions
- How many cuts does the tool miss?
- Measured over ten synthetic stacks, 1,048 pages and 639 documents: 98.3 % of the cuts are found, and none is spurious. Eleven corrections in all, four at worst on one stack.
- What happens when two documents follow one another with no signal?
- That is the limiting case, and it is in the measurement: a whole stack of the corpus is made of documents with no repeated header, no footer and no numbering. The tool finds 96.6 % of the cuts there, and the two it misses take a click each.
- Three hundred pages, does that hold?
- Yes. Reading the text of a 300-page stack takes under a second; reading a fully scanned stack takes 35 seconds with six workers.
- Do the documents keep their quality?
- Yes: pages are copied, not redrawn. A signed page stays what it was, and a text layer stays readable.
The other tools
- Remove blank pages : The empty backs of a scan, set aside and shown, never deleted silently.
- Put a scan back in order : Fronts then backs in reverse: the order is reconstructed, proposed, never imposed.
- Separator sheets : A printed sheet between two documents: the cut becomes certain.
- Identify the documents in a folder : A whole folder read at once: the type of each document, its date, its issuer, the person.
- Identify one document : One document, one answer, and the clues that produced it.