NTH

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

AuthorsYuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang, Qiang Liu, Jiajun Song, Zidun Guo, Xinhan Wang, Handong Zheng, Yang Liu, Dongliang Luo, Zhiyin Ma, Jiarui Zhang, Xiang Bai

July 19, 2026 3 min read
Watch on YouTube
The one-line take

MonkeyOCRv2 turns document images into a specialized visual foundation model that reads text, preserves layout, and powers compact high-performing document AI systems.

Key results

113M
MonkeyDoc v2 scale

Document images used for visual-text pretraining.

17
MonkeyDoc v2 languages

Languages covered by the multilingual pretraining corpus.

58.7%
CRNN recognition baseline

Overall CRNN accuracy with the original visual encoder.

67.3%
CRNN recognition with MonkeyOCRv2

Overall CRNN accuracy after replacing its encoder with MonkeyOCRv2.

83.3%
MDPBench score

Overall score of the 0.7B MonkeyOCRv2-B parsing model.

57.2
Document understanding average

MonkeyOCRv2-B average across eight benchmarks with Qwen3-1.7B.

What the paper found

Researchers at Huazhong University of Science and Technology and Kingsoft Office present MonkeyOCRv2, a document-native visual-text foundation model designed to preserve character strokes, glyph details, and layout information that natural-image encoders such as CLIP, DINO, and SAM often discard. Its MonkeyDoc v2 pretraining corpus contains 113M document images spanning 17 languages, combining expert-labeled real documents with synthetic multilingual pages and cropped elements. The central technique jointly optimizes autoregressive image-to-text generation with pixel-level reconstruction, using mean-squared error and, in document understanding experiments, edge- and distance-aware structural losses. As a transferable backbone, MonkeyOCRv2 raises CRNN text-recognition accuracy from 58.7% to 67.3%, improves formula, detection, tampering, and overlapping-text segmentation results, and lets a 110M UniMERNet-T outperform the 325M UniMERNet-B. Frozen and paired with Qwen3-0.6B, its 0.7B parsing model reaches 83.3% on multilingual MDPBench, up from 80.5% for the 3B dots.mocr, using a roughly 0.1B vision encoder; on OmniDocBench, it outperforms much larger systems including Qwen3-VL-235B and OpenAI’s GPT-5.2, although specialized parsers remain stronger there. Under identical training with Qwen3-1.7B, MonkeyOCRv2-B averages 57.2 across eight document-understanding benchmarks, surpassing OpenVision-B by 13.2 points and outperforming encoders behind models such as Google’s Gemini-3-pro and Anthropic’s Claude in the paper’s broader comparisons. The study argues that document-oriented visual pretraining can serve as an independent foundation for multilingual OCR, parsing, and document intelligence.

Original abstract

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11$\times$ smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis