Fetching the paper…
Reading the bibliography…
We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation.
“Surrey Audio-Visual Expressed Emotion (SAVEE) database”, 2011
Philip Jackson and Sana ul haq · 2011
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research”
Justin Salamon, Christopher Jacoby and Juan Bello · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“ESC: Dataset for Environmental Sound Classification”
Karol. Piczak · 2015
Earlier work this paper cites.
“TUT Database for Acoustic Scene Classification and Sound Event Detection”
Annamaria Mesaros, Toni Heittola and Tuomas Virtanen · 2016
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline”
Hui Bu et al · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization”
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
“Aishell-2: Transforming mandarin asr research into industrial scale”
Jiayu Du, Xingyu Na, Xuechen Liu and Hui Bu · 2018
Earlier work this paper cites.
“The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English”
Steven Livingstone and Frank Russo · 2018
Earlier work this paper cites.
“Meld: A multimodal multi-party dataset for emotion recognition in conversations”
Soujanya Poria et al · 2018
Earlier work this paper cites.
“Multi-modal emotion recognition on iemocap dataset using deep learning”
Samarth Tripathi, Sarthak Tripathi and Homayoon Beigi · 2018
Earlier work this paper cites.
“Common voice: A massively-multilingual speech corpus”
Rosana Ardila et al · 2019
Earlier work this paper cites.
“Audiocaps: Generating captions for audios in the wild”
Chris Kim, Byeongchang Kim, Hyunmin Lee and Gunhee Kim · 2019
Earlier work this paper cites.
“Libritts: A corpus derived from librispeech for text-to-speech”
Heiga Zen et al · 2019
Earlier work this paper cites.
“Vggsound: A large-scale audio-visual dataset”
Honglie Chen, Weidi Xie, Andrea Vedaldi and Andrew Zisserman · 2020
Earlier work this paper cites.
“Clotho: An audio captioning dataset”
Konstantinos Drossos, Samuel Lipping and Tuomas Virtanen · 2020
Earlier work this paper cites.
“MLS: A Large-Scale Multilingual Dataset for Speech Research”
Vineel Pratap et al · 2020
Earlier work this paper cites.
“Aishell-3: A multi-speaker mandarin tts corpus and the baselines”
Yao Shi et al · 2020
Earlier work this paper cites.
“Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio”
Guoguo Chen et al · 2021
Earlier work this paper cites.
“Fsd50k: an open dataset of human-labeled sound events”
Eduardo Fonseca et al · 2021
Earlier work this paper cites.
“What is the ground truth? reliability of multi-annotator data for audio tagging”
Irene Martín-Morató and Annamaria Mesaros · 2021
Earlier work this paper cites.
“Kespeech: An open source speech dataset of mandarin and its eight subdialects”
Zhiyuan Tang et al · 2021
Earlier work this paper cites.
Changhan Wang et al · 2021
Earlier work this paper cites.
Zhifu Gao, Shiliang Zhang, Ian McLoughlin and Zhijie Yan · 2022
Earlier work this paper cites.
“Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition”
Yuan Gong, Jin Yu and James Glass · 2022
Earlier work this paper cites.
“TAU Urban Acoustic Scenes 2022 Mobile, Development dataset”, Zenodo, 2022
T. Heittola et al · 2022
Earlier work this paper cites.
“Cochlscene: Acquisition of acoustic scene data using crowdsourcing”
Il-Young Jeong and Jeongsoo Park · 2022
Earlier work this paper cites.
“Bigvgan: A universal neural vocoder with large-scale training”
Sang-gil Lee et al · 2022
Earlier work this paper cites.
“Learning to answer questions in dynamic audio-visual scenarios”
Guangyao Li et al · 2022
Earlier work this paper cites.
“Clotho-aqa: A crowdsourced dataset for audio question answering”
Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos and Tuomas Virtanen · 2022
Earlier work this paper cites.
“Avqa: A dataset for audio-visual question answering on videos”
Pinci Yang et al · 2022
Cited alongside, same era.
“Open source magicdata-ramc: A rich annotated mandarin conversational (ramc) speech dataset”
Zehui Yang et al · 2022
Cited alongside, same era.
“High fidelity speech enhancement with band-split rnn”
Jianwei Yu et al · 2022
Cited alongside, same era.
“Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition”
Binbin Zhang et al · 2022
Cited alongside, same era.
“Emotional voice conversion: Theory, databases and ESD”
Kun Zhou, Berrak Sisman, Rui Liu and Haizhou Li · 2022
Cited alongside, same era.
“Audiolm: a language modeling approach to audio generation”
Aaron Hurst et al · 2024
Later among the works it cites.
“Libriheavy: A 50,000 hours ASR corpus with punctuation casing and context”
Wei Kang et al · 2024
Later among the works it cites.
“Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities”
Zhifeng Kong et al · 2024
Later among the works it cites.
“T \ \backslash " ulu 3: Pushing frontiers in open language model post-training”
Nathan Lambert et al · 2024
Later among the works it cites.
“Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions”
Jia Li et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zalán Borsos et al · 2023
Cited alongside, same era.
“Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models”
Yunfei Chu et al · 2023
Cited alongside, same era.
“Fleurs: Few-shot learning evaluation of universal representations of speech”
Alexis Conneau et al · 2023
Cited alongside, same era.
“Voicebox: Text-guided multilingual universal speech generation at scale”
Matthew Le et al · 2023
Cited alongside, same era.
“OpenOrca: An Open Dataset of GPT Augmented FLAN Reasoning Traces”
Wing Lian et al · 2023
Cited alongside, same era.
“Music Source Separation With Band-Split RNN”
Yi Luo and Jianwei Yu · 2023
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision”
Alec Radford et al · 2023
Cited alongside, same era.
“Convincing Audio Generation Based on LLM and Speech Tokenization”
Rui-Bo Liu et al · 2024
Later among the works it cites.
“Zero-shot Voice Conversion with Diffusion Transformers”
Songting Liu · 2024
Later among the works it cites.
“Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model benchmark”
Linhan Ma et al · 2024
Later among the works it cites.
“Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research”
Xinhao Mei et al · 2024
Later among the works it cites.
“Mmau: A massive multi-task audio understanding and reasoning benchmark”
S Sakshi et al · 2024
Later among the works it cites.
“Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen llm”
Xiong Wang et al · 2024
Later among the works it cites.
“Mini-omni: Language models can hear, talk while thinking in streaming”
Zhifei Xie and Changqiao Wu · 2024
Later among the works it cites.
“Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing”
Zhangchen Xu et al · 2024
Later among the works it cites.
An Yang et al · 2024
Later among the works it cites.
“Autoprep: An automatic preprocessing framework for in-the-wild speech data”
Jianwei Yu et al · 2024
Later among the works it cites.
“Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot”
Aohan Zeng et al · 2024
Later among the works it cites.
“Omniflatten: An end-to-end gpt model for seamless voice conversation”
Qinglin Zhang et al · 2024
Later among the works it cites.
“MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning”
Hang Zhao et al · 2024
Later among the works it cites.
“Minmo: A multimodal large language model for seamless voice interaction”
Qian Chen et al · 2025
Closest in time.
“OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia”
Xuelong Geng et al · 2025
Closest in time.
“Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation”
Haorui He et al · 2025
Closest in time.
“Step-audio: Unified understanding and generation in intelligent speech interaction”
Ailin Huang et al · 2025
Closest in time.
“MoonCast: High-Quality Zero-Shot Podcast Generation”
Zeqian Ju et al · 2025
Closest in time.
“Baichuan-audio: A unified framework for end-to-end speech interaction”
Tianpeng Li et al · 2025
Closest in time.
“Muon is Scalable for LLM Training”, 2025
Jingyuan Liu et al · 2025
Closest in time.
“Kimi k1.5: Scaling Reinforcement Learning with LLMs”, 2025
Kimi Team et al · 2025
Closest in time.
“Qwen2. 5-omni technical report”
Jin Xu et al · 2025
Closest in time.
“Qwen2.5-Omni Technical Report”, 2025
Jin Xu et al · 2025
Closest in time.
“Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis”
Zhen Ye et al · 2025
Closest in time.