Damaged Document Examples
source collections
Source collections
These are useful places to compare real scanning decisions. I prefer collections that leave enough image evidence to understand the document condition, even when the page is ugly. Some are large public sites; some are smaller collections noted here because the defects are easy to study.
- Internet Archive text collections - uneven capture methods make it useful for comparing scanner beds, cameras, OCR layers, and compression choices.
- Chronicling America newspapers - good for column OCR problems, brittle paper, contrast loss, and page images with uneven exposure.
- County minute book scans - useful for warped ledger pages, punched holes, dark gutter shadows, and handwritten corrections beside typed entries. [local notes only]
- Biodiversity Heritage Library - useful for bound volumes, plates, foldouts, and mixed page sizes captured across many source conditions.
- Photocopied technical circulars - repeated copy generations create gray fog, broken diagrams, and false punctuation in OCR. [index card reference]
- field record digitization project - compact institutional records with visible page wear, edge shadows, handwritten marks, and inconsistent capture quality.
- NYPL Digital Collections - useful for accession cards, correspondence, clippings, and a wide range of paper handling marks.
- Church bulletin archive, 1960s-1980s - folded sheets and low-contrast mimeograph ink show how pale originals fail after thresholding. [offline binder]
- University lab notebook scans - grid paper, pencil annotations, taped inserts, and chemical stains make OCR mostly a finding aid rather than a transcript. [private teaching set]
- Maintenance log collections - narrow columns, carbon copies, abbreviations, and feeder streaks are useful for testing whether a scan is readable before OCR.
None of these should be treated as a model workflow by itself. They are useful because they preserve ordinary flaws that get removed from polished demonstration sets.