Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning
AuthorsTong Ye, Hang Yu, Tengfei Ma, Xuhong Zhang, Jianguo Li, Peng Di, Peiyu Liu, Jianwei Yin, Wenhai Wang
Resources
DOMINO learns the hidden structure of a target coding domain from a few examples and uses it to synthesize better training data for large language models.
Key results
Live Code Generation reference samples used for implicit-domain synthesis
Live Code Execution reference samples used for implicit-domain synthesis
Live Code Generation synthetic inputs generated for comparison
Live Code Execution synthetic inputs generated for comparison
Best improvement reported over strong instruction-tuned backbones on coding benchmarks
DOMINO on Live Code Execution with Qwen2.5-Coder-7B-Instruct
What the paper found
DOMINO, from vivo AI Lab, Ant Group, and Zhejiang University, tackles domain-specific data synthesis when the target domain is only given by reference examples rather than a natural-language specification. The method learns a minimal sufficient domain representation with soft prompt tuning, then adds a contrastive disentanglement objective that separates shared domain structure from sample-specific noise, so the learned prompt generalizes instead of memorizing. On LiveCodeBench’s implicit coding domains, DOMINO is trained on 713 Live Code Generation references and 66 Live Code Execution references, synthesizes 80K and 40K samples respectively, and fine-tuning on its synthetic data lifts Pass@1 by up to 4.63% over strong instruction-tuned backbones. In the main coding results, DOMINO reaches 12.63 Pass@1 on Live Code Generation and 42.59 Pass@1 on Live Code Execution with OPENCODER-8B-Instruct, and 17.31 and 56.48 with Qwen2.5-Coder-7B-Instruct, while also surpassing the instruction-tuned backbone on LiveBench-style instruction following with a 55.39 average score. The paper’s theoretical claim is that disentanglement expands the effective support of the synthetic distribution, and the empirical analyses support that with broader t-SNE spread, stronger robustness across temperatures from 0.2 to 1.0, and better performance as domain soft-token capacity increases from 64 to 512. Overall, DOMINO reframes synthetic data generation as inductive representation learning for domains that are hard to describe explicitly, with Ant Group’s and vivo AI Lab’s implementation showing that self-generated, domain-aligned data can outperform manual prompt design.
Original abstract
Large Language Models have demonstrated remarkable progress in general-purpose capabilities and can achieve strong performance in specific domains through fine-tuning on domain-specific data. However, acquiring high-quality data for target domains remains a significant challenge. Existing data synthesis approaches follow a deductive paradigm, heavily relying on explicit domain descriptions expressed in natural language and careful prompt engineering, limiting their applicability in real-world scenarios where domains are difficult to describe or formally articulate. In this work, we tackle the underexplored problem of domain-specific data synthesis through an inductive paradigm, where the target domain is defined only through a set of reference examples, particularly when domain characteristics are difficult to articulate in natural language. We propose a novel framework, DOMINO, that learns a minimal sufficient domain representation from reference samples and leverages it to guide the generation of domain-aligned synthetic data. DOMINO integrates prompt tuning with a contrastive disentanglement objective to separate domain-level patterns from sample-specific noise, mitigating overfitting while preserving core domain characteristics. Theoretically, we prove that DOMINO expands the support of the synthetic data distribution, ensuring greater diversity. Empirically, on challenging coding benchmarks where domain definitions are implicit, fine-tuning on data synthesized by DOMINO improves Pass@1 accuracy by up to 4.63\% over strong, instruction-tuned backbones, demonstrating its effectiveness and robustness. This work establishes a new paradigm for domain-specific data synthesis, enabling practical and scalable domain adaptation without manual prompt design or natural language domain specifications.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.