MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
AuthorsYuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang, Qiang Liu, Jiajun Song, Zidun Guo, Xinhan Wang, Handong Zheng, Yang Liu, Dongliang Luo, Zhiyin Ma, Jiarui Zhang, Xiang Bai
Resources
MonkeyOCRv2 turns document images into a specialized visual foundation model that reads text, preserves layout, and powers compact high-performing document AI systems.
Key results
Document images used for visual-text pretraining.
Languages covered by the multilingual pretraining corpus.
Overall CRNN accuracy with the original visual encoder.
Overall CRNN accuracy after replacing its encoder with MonkeyOCRv2.
Overall score of the 0.7B MonkeyOCRv2-B parsing model.
MonkeyOCRv2-B average across eight benchmarks with Qwen3-1.7B.
What the paper found
Researchers at Huazhong University of Science and Technology and Kingsoft Office present MonkeyOCRv2, a document-native visual-text foundation model designed to preserve character strokes, glyph details, and layout information that natural-image encoders such as CLIP, DINO, and SAM often discard. Its MonkeyDoc v2 pretraining corpus contains 113M document images spanning 17 languages, combining expert-labeled real documents with synthetic multilingual pages and cropped elements. The central technique jointly optimizes autoregressive image-to-text generation with pixel-level reconstruction, using mean-squared error and, in document understanding experiments, edge- and distance-aware structural losses. As a transferable backbone, MonkeyOCRv2 raises CRNN text-recognition accuracy from 58.7% to 67.3%, improves formula, detection, tampering, and overlapping-text segmentation results, and lets a 110M UniMERNet-T outperform the 325M UniMERNet-B. Frozen and paired with Qwen3-0.6B, its 0.7B parsing model reaches 83.3% on multilingual MDPBench, up from 80.5% for the 3B dots.mocr, using a roughly 0.1B vision encoder; on OmniDocBench, it outperforms much larger systems including Qwen3-VL-235B and OpenAI’s GPT-5.2, although specialized parsers remain stronger there. Under identical training with Qwen3-1.7B, MonkeyOCRv2-B averages 57.2 across eight document-understanding benchmarks, surpassing OpenVision-B by 13.2 points and outperforming encoders behind models such as Google’s Gemini-3-pro and Anthropic’s Claude in the paper’s broader comparisons. The study argues that document-oriented visual pretraining can serve as an independent foundation for multilingual OCR, parsing, and document intelligence.
Original abstract
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11$\times$ smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.