NTH

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

AuthorsDavid Lowry-Duda, Matteo Cargnelutti, Catherine Brobston, Salwa Ismail, Greg Leppert, Amanda Watson, Jonathan Zittrain

August 25, 2026 2 min read
Watch on YouTube
The one-line take

This work turns nearly a million scanned books into a flexible, multilingual research resource by preserving rich annotations instead of making one irreversible cleaning decision.

Key results

217B
IB-HL-ET tokens

The released enriched corpus contains 217B o200k_base tokens.

983,003
Volumes

The corpus covers 983,003 processed volumes.

1.39B
Subtopic paragraphs

Semantic chunking produces 1.39B annotated subtopic paragraphs.

0.97
Endmatter classifier F1

The synthetic-data-trained endmatter classifier achieves an F1 of 0.97.

86.57
IB-HL-ET tokenizability

Overall tokenizability rises to 86.57 from 80.43 in the source OCR.

What the paper found

Institutional Books—Enriched Text introduces a configurable, multilingual pipeline for transforming OCR from historical books without collapsing every editorial decision into one cleaned token stream. Applied across approximately 250 languages, it uses soft and hard Unicode normalization, n-gram language models for context-sensitive dehyphenation, 128-bit simhash with MurmurHash3 for duplicate-page and duplicate-paragraph detection, and a combination of Nupunkt and SaT for sentence segmentation. Semantic chunking adapts TextTiling with CPU-efficient Model2Vec embeddings distilled from BAAI/BGE-M3, while each paragraph receives language, duplicate-cluster, and Qwen/Qwen3-0.6B-Base bits-per-byte annotations. Endmatter classification uses synthetic multilingual examples generated primarily with OpenAI’s gpt-oss-20b, after comparison with Gemma3 and Qwen3, and achieves 0.97 accuracy with an F1 of 0.97. The released IB-HL-ET corpus contains 217B o200k_base tokens across 983,003 volumes and 1.39B subtopic paragraphs. Rather than deleting repeated text or paratext, the pipeline marks it in HTML-like annotations, allowing users to filter endmatter, duplicates, languages, or OCR-quality outliers for retrieval, scholarship, or language-model training. Cleaning raises overall tokenizability from 80.43 in the source OCR to 86.57 in IB-HL-ET, demonstrating that annotation-preserving preprocessing can improve machine usability while retaining provenance and multilingual coverage.

Original abstract

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, deduplicate, restrict by language, and sometimes discard meaningful metadata. Meanwhile, researchers seeking to use IB-HL duplicate effort while performing similar processing and analysis. We describe an approach that we call Enriched Text. Instead of producing a single 'complete' stream of tokens, we normalize the text while preserving metadata through annotations. We separate endmatter, detect per-paragraph language, identify clusters of duplicate paragraphs, and compute per-paragraph bits-per-byte scores. We provide this information through HTML-like annotations layered on top of the text. By parsing these annotations, users can tailor the output to their own needs instead of accepting a global editorial decision on content. The pipeline applies to all $\approx$250 languages in the collection. This report describes this project's goals, implementation, and design rationale. The release includes IB-HL-ET (an enriched-text version of IB-HL containing 217B o200k_base tokens across 983,003 volumes, organized into 1.39B annotated subtopic paragraphs) and the pipeline that produced it. These serve to make the collection easier for machines to parse and for humans to study.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →