NTH

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

AuthorsXingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen

Affiliations[

October 2, 2026 2 min read
Watch on YouTube
The one-line take

A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.

Key results

10M
TextMuSS-10M training samples

Balanced synthetic scene-text data used for multilingual training

229
Supported languages

Languages represented across the 10 target scripts

82.06%
TextMuSS-Bench average accuracy

Average word accuracy achieved by ScriptMoE

1.31%
Improvement over strongest STR baseline

Absolute average-accuracy gain on TextMuSS-Bench

80.89%
CC-OCR F1 with PP-OCRv5

End-to-end multilingual F1 after replacing only the recognizer

41.13M
Activated parameters per image

Parameters used during one ScriptMoE forward pass

What the paper found

This paper presents an all-in-one multilingual scene-text recognizer designed to avoid both per-language OCR systems and heavyweight vision-language models. Its training resource, TextMuSS-10M, contains 10M balanced synthetic scene-text samples spanning 10 scripts and 229 languages, while TextMuSS-Bench evaluates recognition on 10,899 real images. The proposed ScriptMoE combines a shared SVTRv2 visual encoder with an autoregressive decoder whose dense feed-forward layers are replaced by an image-level, script-aware Mixture-of-Experts module: a router selects the top-2 of four script-family experts for every image, and an always-active shared expert preserves cross-script knowledge. A lightweight script-classification loss further guides specialization. ScriptMoE reaches 82.06% average word accuracy on TextMuSS-Bench, improving on the strongest STR baseline by 1.31%, with notable gains for Arabic, Thai, and Tibetan. The model stores 45.85M parameters but activates only 41.13M per image, making it substantially lighter than models such as Qwen3.5-9B. Replacing only the recognizer in Tencent’s PP-OCRv5 pipeline raises CC-OCR multilingual F1 from 65.71% to 80.89%, slightly exceeding Qwen3.5-9B’s 80.73% while retaining a compact OCR-oriented architecture.

Original abstract

Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages. It provides balanced and sufficient supervision where real data is unavailable. Second, we propose ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture. It shares a single visual encoder and replaces the dense decoder with a sparse MoE block, which consists of an image-level router dispatches each image to the top-2 script-aligned experts and a shared expert absorbs cross-script knowledge. Extensive experiments on our assembled TextMuSS-Bench (10 scripts, 10,899 images) show that ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%. On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.

Read the original paper

More in Computer Vision

Browse all 58 papers →
01Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
02Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis