Skip to content
Klarfolder

Scanning a stack of documents

Three settings decide whether the stack will be workable or merely legible to the eye.

300 dots per inch, no less

Below that, printed codes stop being readable. A French 2D-Doc on a tax assessment has a module of about four tenths of a millimetre: it takes at least three pixels per module to decode, and measurement shows that below two and a quarter pixels it is lost.

Above 300 you gain file size and nothing else.

The text layer, if the scanner can produce it

A PDF carrying its text layer reads instantly. A PDF that is only an image has to be read by character recognition, which takes time even with six workers in parallel: thirty-five seconds for three hundred pages, against three tenths of a second.

If the scanner offers “searchable PDF” or “PDF with OCR”, that is it.

What not to set

  • Aggressive auto-straightening: it sometimes cuts an edge, and that is often the edge with the code on it.
  • Heavy compression: it erases small characters, and a compressed code no longer reads.
  • Photo mode on text: it triples the size and adds nothing.

Frequently asked questions

Double-sided?
Scan the fronts, then the backs, in the same order. Putting pages back in order is a known problem, and it solves better than rescanning.
Should documents be separated?
Not necessarily. A stack of three hundred pages can be split into separate documents; that module comes after launch.