NTH

GigaChat Audio: Time-aware Large Audio Language Model

AuthorsAleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin

August 3, 2026 2 min read
Watch on YouTube
The one-line take

GigaChat Audio teaches a large audio-language model to understand long recordings and answer questions with precise timestamps.

Key results

120
Maximum input duration

Maximum recording length supported by GigaChat Audio.

14k
Filtered YODAS2 audio

Hours retained after language and silence-ratio filtering.

53.8
Long-form grounding at 60-second anchors

mIoU on 20–40-minute recordings.

14.2
Long-form grounding without anchors

mIoU after removing inter-timing markers on 20–40-minute recordings.

65.2
Grounding at 7-second anchors

mIoU on 20–40-minute recordings with more frequent temporal anchors.

3
Timing error at 60-second anchors

Median absolute midpoint error in seconds.

What the paper found

Researchers at SaluteDevices present GigaChat Audio, an open-weight time-aware Audio LLM built on a 10B-A1.8B mixture-of-experts language model that processes recordings up to 120 minutes and generates answers, descriptions, and summaries anchored to explicit timestamps. Its central technique interleaves continuous audio tokens with periodic temporal anchors, using a cascaded synthetic-supervision pipeline: WhisperX aligns YODAS2 transcripts, GPT-OSS-120B generates timestamped questions and answers from roughly 10-minute slices, and a global verifier filters inconsistencies. On 20–40-minute temporal grounding, the model reaches 53.8 mIoU with anchors every 60 seconds, while removing anchors drops performance to 14.2 mIoU; increasing anchor frequency to every 7 seconds raises accuracy to 65.2 mIoU. At 60-second spacing, midpoint timing error is 3 seconds, showing that sparse anchors can support fine-grained localization through interpolation. The training corpus retains 14k hours after filtering from multilingual YODAS2 audio and mixes recordings across duration ranges, because short-only training fails to extrapolate to long recordings and long-only training harms short-audio performance. Compared with systems such as OpenAI’s GPT-OSS-120B used for data generation, Qwen3-Omni, and Google’s Gemini 3 Flash, the work emphasizes verifiable long-form temporal grounding rather than general audio question answering, and releases model weights plus a 10k+ hour temporal dataset.

Original abstract

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis