GigaChat Audio: Time-aware Large Audio Language Model
AuthorsAleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin
Resources
GigaChat Audio teaches a large audio-language model to understand long recordings and answer questions with precise timestamps.
Key results
Maximum recording length supported by GigaChat Audio.
Hours retained after language and silence-ratio filtering.
mIoU on 20–40-minute recordings.
mIoU after removing inter-timing markers on 20–40-minute recordings.
mIoU on 20–40-minute recordings with more frequent temporal anchors.
Median absolute midpoint error in seconds.
What the paper found
Researchers at SaluteDevices present GigaChat Audio, an open-weight time-aware Audio LLM built on a 10B-A1.8B mixture-of-experts language model that processes recordings up to 120 minutes and generates answers, descriptions, and summaries anchored to explicit timestamps. Its central technique interleaves continuous audio tokens with periodic temporal anchors, using a cascaded synthetic-supervision pipeline: WhisperX aligns YODAS2 transcripts, GPT-OSS-120B generates timestamped questions and answers from roughly 10-minute slices, and a global verifier filters inconsistencies. On 20–40-minute temporal grounding, the model reaches 53.8 mIoU with anchors every 60 seconds, while removing anchors drops performance to 14.2 mIoU; increasing anchor frequency to every 7 seconds raises accuracy to 65.2 mIoU. At 60-second spacing, midpoint timing error is 3 seconds, showing that sparse anchors can support fine-grained localization through interpolation. The training corpus retains 14k hours after filtering from multilingual YODAS2 audio and mixes recordings across duration ranges, because short-only training fails to extrapolate to long recordings and long-only training harms short-audio performance. Compared with systems such as OpenAI’s GPT-OSS-120B used for data generation, Qwen3-Omni, and Google’s Gemini 3 Flash, the work emphasizes verifiable long-form temporal grounding rather than general audio question answering, and releases model weights plus a 10k+ hour temporal dataset.
Original abstract
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.