NTH

Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning

AuthorsTong Ye, Hang Yu, Tengfei Ma, Xuhong Zhang, Jianguo Li, Peng Di, Peiyu Liu, Jianwei Yin, Wenhai Wang

June 29, 2026 3 min read
Watch on YouTube
The one-line take

DOMINO learns the hidden structure of a target coding domain from a few examples and uses it to synthesize better training data for large language models.

Key results

713
LiveCodeBench reference set

Live Code Generation reference samples used for implicit-domain synthesis

66
LiveCodeBench reference set

Live Code Execution reference samples used for implicit-domain synthesis

80K
Synthetic generation scale

Live Code Generation synthetic inputs generated for comparison

40K
Synthetic generation scale

Live Code Execution synthetic inputs generated for comparison

4.63%
Pass@1 gain

Best improvement reported over strong instruction-tuned backbones on coding benchmarks

56.48
Best Pass@1

DOMINO on Live Code Execution with Qwen2.5-Coder-7B-Instruct

What the paper found

DOMINO, from vivo AI Lab, Ant Group, and Zhejiang University, tackles domain-specific data synthesis when the target domain is only given by reference examples rather than a natural-language specification. The method learns a minimal sufficient domain representation with soft prompt tuning, then adds a contrastive disentanglement objective that separates shared domain structure from sample-specific noise, so the learned prompt generalizes instead of memorizing. On LiveCodeBench’s implicit coding domains, DOMINO is trained on 713 Live Code Generation references and 66 Live Code Execution references, synthesizes 80K and 40K samples respectively, and fine-tuning on its synthetic data lifts Pass@1 by up to 4.63% over strong instruction-tuned backbones. In the main coding results, DOMINO reaches 12.63 Pass@1 on Live Code Generation and 42.59 Pass@1 on Live Code Execution with OPENCODER-8B-Instruct, and 17.31 and 56.48 with Qwen2.5-Coder-7B-Instruct, while also surpassing the instruction-tuned backbone on LiveBench-style instruction following with a 55.39 average score. The paper’s theoretical claim is that disentanglement expands the effective support of the synthetic distribution, and the empirical analyses support that with broader t-SNE spread, stronger robustness across temperatures from 0.2 to 1.0, and better performance as domain soft-token capacity increases from 64 to 512. Overall, DOMINO reframes synthetic data generation as inductive representation learning for domains that are hard to describe explicitly, with Ant Group’s and vivo AI Lab’s implementation showing that self-generated, domain-aligned data can outperform manual prompt design.

Original abstract

Large Language Models have demonstrated remarkable progress in general-purpose capabilities and can achieve strong performance in specific domains through fine-tuning on domain-specific data. However, acquiring high-quality data for target domains remains a significant challenge. Existing data synthesis approaches follow a deductive paradigm, heavily relying on explicit domain descriptions expressed in natural language and careful prompt engineering, limiting their applicability in real-world scenarios where domains are difficult to describe or formally articulate. In this work, we tackle the underexplored problem of domain-specific data synthesis through an inductive paradigm, where the target domain is defined only through a set of reference examples, particularly when domain characteristics are difficult to articulate in natural language. We propose a novel framework, DOMINO, that learns a minimal sufficient domain representation from reference samples and leverages it to guide the generation of domain-aligned synthetic data. DOMINO integrates prompt tuning with a contrastive disentanglement objective to separate domain-level patterns from sample-specific noise, mitigating overfitting while preserving core domain characteristics. Theoretically, we prove that DOMINO expands the support of the synthetic data distribution, ensuring greater diversity. Empirically, on challenging coding benchmarks where domain definitions are implicit, fine-tuning on data synthesized by DOMINO improves Pass@1 accuracy by up to 4.63\% over strong, instruction-tuned backbones, demonstrating its effectiveness and robustness. This work establishes a new paradigm for domain-specific data synthesis, enabling practical and scalable domain adaptation without manual prompt design or natural language domain specifications.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis