NTH

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

AuthorsYuyue Wang, Xihua Wang, Xin Cheng, Yijing Chen, Ruihua Song

June 2, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes a new way to generate combined speech and sound directly from natural language using an LLM-based planner and a new benchmark for composite audio.

Key results

1.27M
Training pool

Total training data pool across composite, sound, and speech scenarios.

4500
PlanAudio-Bench held-out composite examples

Held-out composite examples used to form PlanAudio-Bench from AudioSet.

371K
Composite training split

Composite clips included in the 1.27M training pool.

8.52
FAD on PlanAudio-Bench

PlanAudio’s composite-scenario Fr’echet Audio Distance, improving over pipeline and unified baselines.

0.41
WER on PlanAudio-Bench

PlanAudio’s composite-scenario Word Error Rate, reported for speech-related evaluation.

1.5B
Qwen2.5 initialization size

PlanAudio is initialized from Qwen2.5-1.5B as the autoregressive LLM backbone.

What the paper found

Renmin University of China’s PlanAudio reframes audio generation as Free-Form-Text-Prompt-to-Unified-Audio synthesis, where a single unconstrained prompt must produce speech, sound, or tightly interleaved composites without external text rewriting. The key novelty is a unified autoregressive LLM backbone initialized from Qwen2.5-1.5B that bypasses a separate text encoder and inserts a semantic latent Chain-of-Thought stage: the model first predicts a compact continuous planning sequence z, supervised by segment-level embeddings from Audio Flamingo 3, then generates hierarchical AudioCraft codec tokens conditioned on both the prompt and z. To evaluate this setting, the authors build PlanAudio-Bench from AudioSet with decoupled annotations via Whisper and Gemini-2.5 Pro, using 4,500 held-out composite examples and a 1.27M-example training pool spanning 371k composite, 451k sound, and 354k speech clips. Across composite benchmarks, PlanAudio improves substantially over pipeline and unified baselines such as VoiceLDM and AudioLDM2 combinations, reducing FAD from 22.9–25.2 down to 8.52 and lowering WER to 0.41 while achieving the best human ratings for quality, temporal correctness, semantic alignment, and authenticity. Ablations show that semantic latent CoT outperforms explicit natural-language CoT and acoustic-token CoT, and that a constant multi-scenario curriculum is crucial to avoid catastrophic forgetting.

Original abstract

Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis