Boston Public Library and the Institutional Data Initiative at Harvard Law School Library have released a data set derived from public domain newspaper collections available on DigitalCommonwealth.org.
Newspapers represent one of the richest available primary sources for historical research — capturing not just major news, but births and deaths, iconic sports victories (or crushing defeats), political debates, advertisements, fashion, entertainment listings, and other content. But a major problem is that historical newspapers are a notoriously difficult challenge for optical character recognition (OCR) — the part of the digitization process that converts page images into searchable text. Dense columns, decorative typefaces, and the ravages of age mean the extracted text is often garbled or incomplete.
To address this challenge, IDI developed a new approach using machine learning and AI to isolate individual articles and segments within newspaper pages, then ran text extraction on each one separately — a more targeted approach than processing whole pages at once. Across the full dataset, that process produced billions of words of high-quality extracted text.
Segmented page from the Springfield Weekly Republican, August 19, 1915:

A single segment from The Evening Union (Springfield, Mass.), December 28, 1910:

Traditional OCR (extraneous punctuation, misspelled words, gibberish, missing content, etc.):
| —————; .
| Funeral of Mrs. E. M. Hannan.
The funeral of Mrs. Eva M. Han-
nan was held from the home of her
sister in law, Mrs. Clifford a eyo,
Oscar
yesterday at 2 o’clock, Rev,
Enhanced OCR:
Funeral of Mrs. E. M. Hannan.
The funeral of Mrs. Eva M. Hannan was held from the home of her sister in law, Mrs. Clifford Prevost, yesterday at 2 o’clock, Rev. C. Oscar Ford officiating. The body was taken to Hazardville, Conn., by special car, where burial will take place in the family lot.
The dataset includes not just higher-accuracy text, but also what type of content it is — such as news articles, advertisements, birth notices, literary works, and illustrations — plus the names of people, places, and organizations mentioned.
The newspapers in the dataset span Massachusetts broadly, with titles from Boston neighborhoods like Charlestown, Dorchester, and Roxbury alongside papers from Worcester, Springfield, Salem, New Bedford, and beyond. The dataset also reflects the diversity of the region’s communities: while most content is in English, it also includes newspapers in Yiddish, German, Swedish, and French.
This data set includes 1,473,635 newspaper scans from issues published between 1795 and 1930. This process has produced 83,147,041 individual crops segmented from those scans, over 16 billion o200k_base tokens of VLM OCR text, as well as bounding box coordinates, raw OCR, text analysis, crop type classification, language detection, named-entity recognition, subject classification, reading order detection, and text + image vector embeddings for each segment.
The publication of this data set will be of significant use for computational linguistics, AI model training, and historical research. The open-source processing pipeline built for this project is available for any library or archive to adapt for its own newspaper collections. Designed from the outset to run on workstation-level hardware, it can help unlock millions of pages of historical newspapers with greater accuracy than costly commercial solutions, at a fraction of the cost of commercial OCR software.
The pipeline code, dataset, and full technical documentation are all openly available:
