Reference guide · cited sources
Scanning Documents: TIFF, PDF, or PNG?
Scanning decisions are hard to undo, because the physical original often becomes inaccessible afterwards. Getting the resolution and format right the first time costs a few minutes; getting them wrong can mean rescanning a box of paper you no longer have.
By Prashantkumar Kishanrao Sundge · Published · Reviewed
What this guide establishes
- For a document people will read and share, scan to PDF. It is a multi-page container, universally readable, and can hold searchable text.
- For an archival master of something irreplaceable, scan to TIFF. It is lossless, stable for decades, and holds high bit depths.
- PNG is for single images — a photograph, a diagram, one page you need as a picture. It is not a document format.
- 300 PPI is the working default for text. Go higher for fine print or anything you may need to enlarge; higher still for photographs.
- Text recognition makes a scan searchable. Run it at scan time, because rescanning later is far more work.
Decide what the scan is for
Almost every scanning mistake comes from not separating two different goals. One is producing a readable, shareable copy of a document. The other is producing a faithful preservation copy of an object. They want different formats, different resolutions, and sometimes different equipment.
A readable copy needs to be convenient: one file, openable anywhere, searchable if possible, small enough to email. A preservation copy needs to be faithful: lossless, high resolution, colour-accurate, and stored in a format likely to remain readable in twenty years. Optimising for one at the expense of the other is how people end up with either an unusable 400 MB file or an archive of compressed JPEGs that cannot be re-derived.
The resolution is to do both when the original justifies it: scan once at archival quality, keep that as the master, and generate readable derivatives from it. The scan is the expensive step; making a second file from the first is trivial.
PDF for documents people will use
PDF is a document container. It holds multiple pages natively, keeps them in one file, is readable on essentially every device without special software, and prints predictably. For anything that is fundamentally a document — a contract, a report, a set of records, a manual — it is the right answer and the alternatives are worse.
This matters more than people expect when the alternative is images. A ten-page scan saved as ten PNG files is ten things to attach, ten things to keep in order, and ten things for a recipient to open. The same scan as one PDF is one file that opens to page one and scrolls.
PDF/A is a constrained variant intended specifically for long-term preservation. It forbids features that make a file dependent on external resources — linked fonts, external content, some forms of encryption — so the document remains self-contained and renderable far into the future. Where an institution requires an archival PDF, this is usually what they mean.
PDF also carries a searchable text layer, which is what makes the difference between a scan you can find things in and a scan you have to read through.
TIFF for archival masters
TIFF is lossless, supports high bit depths and CMYK, carries extensive metadata, and has been stable for decades. Those properties are exactly what an archival master needs and are largely irrelevant for a file someone is going to email.
Cultural heritage and records-management practice has long treated TIFF as a preservation format for precisely these reasons: nothing degrades on repeated access, the format is thoroughly documented, and the enormous installed base makes it a safe bet for future readability.
The cost is size. A high-resolution lossless colour scan of a single page can run to tens or hundreds of megabytes, and browsers will not display it. That is acceptable for a master you access rarely and unacceptable for daily use — which is the argument for deriving a PDF from it rather than choosing between them.
Note that a multi-page TIFF holds several complete images in one file. Converting it to a single-image format such as PNG necessarily keeps only one frame; PixelConvert takes the first and reports that it has done so.
Choosing a resolution
For ordinary printed text, 300 PPI is the working default. It resolves normal type comfortably, it is what text recognition engines generally expect, and it keeps file sizes manageable.
Go higher when the original has fine detail: small print, footnotes, handwriting, engravings, or anything you may later want to enlarge or examine closely. 600 PPI is the usual next step and roughly quadruples the file size.
Go higher still for photographs and artwork, where the limit is the original's own detail rather than the size of letterforms. For a small photographic print you intend to reproduce larger, work out the resolution from the output size you need: scanning a 4×6 print that will be reproduced at 8×12 means doubling the linear dimensions, so 600 PPI rather than 300.
Do not scan text in colour by default. Greyscale is usually sufficient for black-on-white text and produces substantially smaller files; pure black-and-white is smaller again but loses the greys that make scanned text legible and is unforgiving of uneven originals. Use colour where colour carries information — letterheads, annotations in ink, stamps, illustrations.
And do not scan low because storage feels tight. Storage is cheap and rescanning is not. Where the original is irreplaceable, err upward.
Text recognition
A scan is a picture of text. Optical character recognition reads that picture and produces actual text, which can then be searched, copied, and indexed. Most scanner software and many PDF tools can apply it during or after scanning.
The usual output keeps the scanned image visible and stores the recognised text invisibly behind it. The document still looks exactly like the original, and a search finds words inside it. This is almost always what you want — it adds searchability without altering appearance and without requiring you to trust the recognition to be perfect.
Accuracy depends heavily on input quality. Clean, straight, well-lit scans of ordinary printed type recognise very well. Skewed pages, poor contrast, unusual typefaces, handwriting, and low resolution all degrade it sharply. Straightening a page before recognition helps more than almost anything else.
Run recognition at scan time. Doing it later means reopening the files, reprocessing, and re-saving — and if the original has gone back in a box or a filing cabinet, any problem you discover is expensive to fix.
Practical workflow
For everyday documents — a signed form, a receipt, a letter to send to someone — scan straight to PDF at 300 PPI in greyscale, with text recognition on. One file, small, searchable, sends anywhere. There is no reason to complicate this.
For records you need to keep but rarely open, the same, with the addition of consistent file naming and a sensible folder structure. The organisation matters more than the format at this point; a perfectly scanned document nobody can find is not much use.
For irreplaceable material — family documents, photographs, anything historical or legal — scan once at high resolution to lossless TIFF, store that as the master, and derive a PDF for use. Keep both. This is the only case where the extra effort is clearly justified, and it is also the case where getting it wrong matters most.
Whatever you choose, verify before disposing of anything. Open the file, check every page is present and the right way up, confirm the text is legible at normal zoom, and check that recognition worked if you enabled it. A scan that failed silently is discovered at exactly the wrong moment.
How this guide was written
This is a reference guide. It explains published specifications, platform rules, and format behavior, and every factual claim is attributed to the sources listed below. It does not report an in-house PixelConvert measurement, and no result here should be read as one.
Read how PixelConvert separates measured and reference guides →Limitations
- This guide summarises widely used scanning and preservation practice with citations; it reports no in-house PixelConvert measurement.
- Resolution recommendations are conventions for common cases, not measured thresholds; specific institutional or legal requirements override them.
- PixelConvert does not produce PDF output and does not perform text recognition. It converts the first frame of a TIFF to PNG.
Frequently asked questions
- Should I scan documents to PDF or images?
- PDF for anything that is a document. It holds multiple pages in one file, opens everywhere, prints predictably, and can carry a searchable text layer. A multi-page scan saved as separate images is harder to use in every respect.
- What resolution should I scan at?
- 300 PPI for ordinary printed text. 600 for fine print, handwriting, or anything you may enlarge. For photographs, work backwards from the size you intend to reproduce at rather than from the original's size.
- Why would I use TIFF instead of PDF?
- For an archival master of something irreplaceable. TIFF is lossless, supports high bit depths, and is a long-established preservation format. It is large and browsers will not display it, so derive a PDF from it for actual use.
- Should I scan in colour or greyscale?
- Greyscale for black-on-white text — it is smaller and entirely sufficient. Colour where colour carries information, such as letterheads, ink annotations, stamps, or illustrations.
- What does text recognition actually do?
- It reads the picture of the text and stores the recognised words invisibly behind the scanned image. The document looks unchanged but becomes searchable and copyable. Run it at scan time, since doing it later means reprocessing everything.
- Can I convert a multi-page TIFF to PNG?
- Only the first page. A PNG holds a single image, so the remaining frames have nowhere to go. If you need all the pages, convert to PDF or use software that exports each frame separately.