Sizing a Digitised Document Collection in KB and MB
Scanning software almost always reports its work one page at a time, and it reports it in kilobytes. A batch log will tell you that page 41 came out at 348 KB and page 42 at 366 KB, but nobody catalogues a collection page by page. The number that matters to an archivist is the megabyte figure for a finished title, a shelf, or a whole accession — and getting there means moving between the two units constantly while you plan a digitisation run.
Per Page Before Per Title
Two Copies of Everything
Multiplied by the Catalogue
The two units also arrive from different directions. Scanner logs, OCR reports and file listings speak kilobytes; vendor quotes, ingest forms and reader-facing download notices speak megabytes. Converting between them is the small routine step that keeps a digitisation estimate from drifting by a factor of a thousand.
Turning Per-Page KB Figures Into Per-Title MB
The converter has two editable fields and reacts as you type, so a whole batch report can be worked through without touching a calculator.
Total the Pages First
Multiply the average page size by the page count while you are still in kilobytes — 412 pages at 340 KB is 140 080 KB. Working in one unit until the very end avoids compounding rounding across a long title list.
Type the Kilobyte Total
Enter it in the left field. Spaces are ignored and a comma is read the same as a dot, so a figure pasted straight out of a spreadsheet or a European-formatted log is accepted as it stands.
Copy the Megabyte Result
The megabyte value appears immediately on the other side. The copy button on that field puts the bare number on the clipboard — no unit, no spacing — which is exactly what a catalogue field or a collection spreadsheet wants.
Change Either Unit When Needed
Both sides carry a searchable unit list. When an accession total outgrows megabytes, pick a larger unit on the right and the same input answers the bigger question without retyping.
Going the other way: the swap button reverses the direction to MB → KB, which is the move you need when a supplier quotes a 4.5 MB average per volume and you want to know what that implies for a single page. Either field can also be typed into directly, so entering a megabyte figure on the right produces the kilobyte equivalent on the left without swapping at all.
Document and E-Book Sizes From Page to Title
The figures below are the ones a digitisation plan actually turns on: what a single scanned page costs at different quality settings, and what an entire text runs to once it is packaged as an e-book. Format matters far more than page count — the same 300-page book can be under a megabyte or well past seventy.
| Item | Kilobytes | Megabytes |
|---|---|---|
| Bitonal 300 dpi TIFF, one letter page | 60 KB | 0.06 MB |
| OCR text layer added to a 200-page scan | 120 KB | 0.12 MB |
| Grayscale 300 dpi JPEG access page | 350 KB | 0.35 MB |
| Reflowable EPUB, 300-page text novel | 800 KB | 0.8 MB |
| Born-digital PDF of the same text | 1 500 KB | 1.5 MB |
| Uncompressed 24-bit 300 dpi TIFF master page | 25 245 KB | 25.245 MB |
| 200-page volume as JPEG access copy | 70 000 KB | 70 MB |
Read the first and the last rows together: a bitonal page and a full colour master page differ by a factor of roughly four hundred. That single choice, taken once at the start of a project, sets the storage bill for everything that follows.
Built for Batch Logs
Paste a kilobyte total straight from a scanner report or a ls listing; stray spaces and a comma decimal mark are both tolerated, so figures rarely need cleaning up first.
Master and Access in One Pass
Both panels are editable and both update live, so you can compare the master stream and the derivative stream by typing alternately into the two sides.
Catalogue-Ready Output
Copying a result yields the bare number with no unit attached — the form a spreadsheet column or a metadata field expects.
Nothing Leaves the Reading Room
Every calculation runs in the browser after the page loads, so unpublished accession figures are never sent anywhere.
Digitisation Sizing Questions
Why is an EPUB of a novel so much smaller than a scanned PDF of the same book?
An EPUB is a compressed bundle of markup and styling — the text itself, stored once, with the reader's device doing the typesetting. A scanned PDF stores a photograph of every page, so it carries several hundred images whether the pages hold a paragraph or a full plate. That is why a 300-page text can sit around 800 KB as an EPUB and reach 70 000 KB as a scan. A born-digital PDF, where the text is real text rather than pictures, lands much closer to the EPUB.
How much does adding an OCR text layer increase a scanned file?
Far less than most people expect. The recognised text is stored as an invisible layer behind the images, and plain text is tiny once compressed: the couple of thousand characters on a typical page shrink to well under a kilobyte, so a 200-page volume gains on the order of 120 KB — around 0.12 MB. What can inflate a file during OCR is the processing step re-saving the page images at different settings, not the text layer itself. If a scan grows noticeably after recognition, look at the image compression options rather than blaming the OCR.
What scanning resolution should I choose if file size matters?
File size climbs with the square of the resolution, so doubling from 300 to 600 dpi does not double the page — it roughly quadruples it. For printed text, 300 dpi is the usual working point: it is enough for reliable character recognition and keeps a colour master page near 25 MB rather than a hundred. Push higher only where the material justifies it — fine engravings, tight manuscript hands, or anything where a reader will want to magnify beyond the page as printed.
How do I estimate storage for a 10 000-title catalogue?
Take a real sample rather than a guess: scan thirty representative titles, average the per-title kilobyte totals, then convert once. At an average of 4 500 KB, or 4.5 MB, per title, ten thousand titles come to 45 000 MB. Do the sum separately for masters and for access copies, because the masters will dominate. Then add headroom for the derivatives you have not thought of yet — thumbnails, page-turner tiles, and the second copy your preservation policy requires.
Should I keep both a preservation master and an access copy?
Yes, and the size gap is the reason the two-copy model works. The master is captured once at full quality and then left alone; the access copy is a lighter derivative that readers actually download, and it can be regenerated at any time from the master. A page held at 25 245 KB as an uncompressed master might be served at 350 KB, so the delivery side of the collection costs a fraction of the archival side. Keep the master's figures out of your public-facing bandwidth estimates entirely.
No comments yet. Be the first to comment!