All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
AuthorsXingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
Affiliations[
Resources
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
Key results
Balanced synthetic scene-text data used for multilingual training
Languages represented across the 10 target scripts
Average word accuracy achieved by ScriptMoE
Absolute average-accuracy gain on TextMuSS-Bench
End-to-end multilingual F1 after replacing only the recognizer
Parameters used during one ScriptMoE forward pass
What the paper found
This paper presents an all-in-one multilingual scene-text recognizer designed to avoid both per-language OCR systems and heavyweight vision-language models. Its training resource, TextMuSS-10M, contains 10M balanced synthetic scene-text samples spanning 10 scripts and 229 languages, while TextMuSS-Bench evaluates recognition on 10,899 real images. The proposed ScriptMoE combines a shared SVTRv2 visual encoder with an autoregressive decoder whose dense feed-forward layers are replaced by an image-level, script-aware Mixture-of-Experts module: a router selects the top-2 of four script-family experts for every image, and an always-active shared expert preserves cross-script knowledge. A lightweight script-classification loss further guides specialization. ScriptMoE reaches 82.06% average word accuracy on TextMuSS-Bench, improving on the strongest STR baseline by 1.31%, with notable gains for Arabic, Thai, and Tibetan. The model stores 45.85M parameters but activates only 41.13M per image, making it substantially lighter than models such as Qwen3.5-9B. Replacing only the recognizer in Tencent’s PP-OCRv5 pipeline raises CC-OCR multilingual F1 from 65.71% to 80.89%, slightly exceeding Qwen3.5-9B’s 80.73% while retaining a compact OCR-oriented architecture.
Original abstract
Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages. It provides balanced and sufficient supervision where real data is unavailable. Second, we propose ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture. It shares a single visual encoder and replaces the dense decoder with a sparse MoE block, which consists of an image-level router dispatches each image to the top-2 script-aligned experts and a shared expert absorbs cross-script knowledge. Extensive experiments on our assembled TextMuSS-Bench (10 scripts, 10,899 images) show that ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%. On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.
Read the original paperMore in Computer Vision
Browse all 58 papers →DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.
A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models
Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu
The study shows that video models may retain the right physical understanding even after generating the wrong motion—and that targeted internal writes can bring that knowledge back.