Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale
AuthorsDavid Lowry-Duda, Matteo Cargnelutti, Catherine Brobston, Salwa Ismail, Greg Leppert, Amanda Watson, Jonathan Zittrain
Resources
This work turns nearly a million scanned books into a flexible, multilingual research resource by preserving rich annotations instead of making one irreversible cleaning decision.
Key results
The released enriched corpus contains 217B o200k_base tokens.
The corpus covers 983,003 processed volumes.
Semantic chunking produces 1.39B annotated subtopic paragraphs.
The synthetic-data-trained endmatter classifier achieves an F1 of 0.97.
Overall tokenizability rises to 86.57 from 80.43 in the source OCR.
What the paper found
Institutional Books—Enriched Text introduces a configurable, multilingual pipeline for transforming OCR from historical books without collapsing every editorial decision into one cleaned token stream. Applied across approximately 250 languages, it uses soft and hard Unicode normalization, n-gram language models for context-sensitive dehyphenation, 128-bit simhash with MurmurHash3 for duplicate-page and duplicate-paragraph detection, and a combination of Nupunkt and SaT for sentence segmentation. Semantic chunking adapts TextTiling with CPU-efficient Model2Vec embeddings distilled from BAAI/BGE-M3, while each paragraph receives language, duplicate-cluster, and Qwen/Qwen3-0.6B-Base bits-per-byte annotations. Endmatter classification uses synthetic multilingual examples generated primarily with OpenAI’s gpt-oss-20b, after comparison with Gemma3 and Qwen3, and achieves 0.97 accuracy with an F1 of 0.97. The released IB-HL-ET corpus contains 217B o200k_base tokens across 983,003 volumes and 1.39B subtopic paragraphs. Rather than deleting repeated text or paratext, the pipeline marks it in HTML-like annotations, allowing users to filter endmatter, duplicates, languages, or OCR-quality outliers for retrieval, scholarship, or language-model training. Cleaning raises overall tokenizability from 80.43 in the source OCR to 86.57 in IB-HL-ET, demonstrating that annotation-preserving preprocessing can improve machine usability while retaining provenance and multilingual coverage.
Original abstract
Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, deduplicate, restrict by language, and sometimes discard meaningful metadata. Meanwhile, researchers seeking to use IB-HL duplicate effort while performing similar processing and analysis. We describe an approach that we call Enriched Text. Instead of producing a single 'complete' stream of tokens, we normalize the text while preserving metadata through annotations. We separate endmatter, detect per-paragraph language, identify clusters of duplicate paragraphs, and compute per-paragraph bits-per-byte scores. We provide this information through HTML-like annotations layered on top of the text. By parsing these annotations, users can tailor the output to their own needs instead of accepting a global editorial decision on content. The pipeline applies to all $\approx$250 languages in the collection. This report describes this project's goals, implementation, and design rationale. The release includes IB-HL-ET (an enriched-text version of IB-HL containing 217B o200k_base tokens across 983,003 volumes, organized into 1.39B annotated subtopic paragraphs) and the pipeline that produced it. These serve to make the collection easier for machines to parse and for humans to study.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.