Fetching the paper…
Reading the bibliography…
This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
“GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio”
Guoguo Chen et al · 1965
Earlier work this paper cites.
“wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations”, 2020
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed and Michael Auli · 2006
Earlier work this paper cites.
“CoVoST 2 and Massively Multilingual Speech-to-Text Translation”, 2020
Changhan Wang, Anne Wu and Juan Pino · 2007
Earlier work this paper cites.
“HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis”, 2020
Jungil Kong, Jaehyeon Kim and Jaekyoung Bae · 2010
Earlier work this paper cites.
“Verbmobil: foundations of speech-to-speech translation”
Wolfgang Wahlster · 2013
Earlier work this paper cites.
“Google’s neural machine translation system: Bridging the gap between human and machine translation”
Yonghui Wu et al · 2016
Earlier work this paper cites.
“Audio Set: An ontology and human-labeled dataset for audio events”
Jort. Gemmeke et al · 2017
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis”
Yuxuan Wang et al · 2017
Earlier work this paper cites.
“Deep voice 3: 2000-speaker neural text-to-speech”
Wei Ping et al · 2018
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions”
Jonathan Shen et al · 2018
Earlier work this paper cites.
“Direct speech-to-speech translation with a sequence-to-sequence model”
Ye Jia et al · 2019
Earlier work this paper cites.
“Audiocaps: Generating captions for audios in the wild”
Chris Kim, Byeongchang Kim, Hyunmin Lee and Gunhee Kim · 2019
Earlier work this paper cites.
“Fastspeech 2: Fast and high-quality end-to-end text to speech”
Yi Ren et al · 2020
Earlier work this paper cites.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units”, 2021
Wei-Ning Hsu et al · 2021
Earlier work this paper cites.
“Textless speech-to-speech translation on real data”
Ann Lee et al · 2021
Earlier work this paper cites.
“Finetuned language models are zero-shot learners”
Jason Wei et al · 2021
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec”
Neil Zeghidour et al · 2021
Earlier work this paper cites.
“BEATs: Audio Pre-Training with Acoustic Tokenizers”, 2022
Sanyuan Chen et al · 2022
Earlier work this paper cites.
“WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing”
Sanyuan Chen et al · 2022
Earlier work this paper cites.
“High fidelity neural audio compression”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve and Yossi Adi · 2022
Earlier work this paper cites.
“Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition”
Yuan Gong, Jin Yu and James Glass · 2022
Earlier work this paper cites.
“CochlScene: Acquisition of acoustic scene data using crowdsourcing”
Il-Young Jeong and Jeongsoo Park · 2022
Earlier work this paper cites.
“CVSS Corpus and Massively Multilingual Speech-to-Speech Translation”, 2022
Ye Jia, Michelle Ramanovich, Quan Wang and Heiga Zen · 2022
Earlier work this paper cites.
“Translatotron 2: High-quality direct speech-to-speech translation with voice preservation”
Ye Jia, Michelle Ramanovich, Tal Remez and Roi Pomerantz · 2022
Earlier work this paper cites.
“Introducing ChatGPT” Accessed: 2025-07-11, 2022
OpenAI · 2022
Earlier work this paper cites.
“WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition”, 2022
Binbin Zhang et al · 2022
Earlier work this paper cites.
“PaLM 2 Technical Report”, 2023
Rohan Anil et al · 2023
Earlier work this paper cites.
Jinze Bai et al · 2023
Earlier work this paper cites.
“Better speech synthesis through scaling”, 2023
James Betker · 2023
Cited alongside, same era.
“Audiolm: a language modeling approach to audio generation”
Zalán Borsos et al · 2023
Cited alongside, same era.
“Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models”
Yunfei Chu et al · 2023
Cited alongside, same era.
“Pengi: An audio language model for audio tasks”
Soham Deshmukh, Benjamin Elizalde, Rita Singh and Huaming Wang · 2023
Cited alongside, same era.
“Joint audio and speech understanding”
Yuan Gong et al · 2023
Cited alongside, same era.
“Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling”
Shengpeng Ji et al · 2024
Later among the works it cites.
“Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities”, 2024
Zhifeng Kong et al · 2024
Later among the works it cites.
“Transvip: Speech to speech translation system with voice and isochrony preservation”
Chenyang Le et al · 2024
Later among the works it cites.
Guan-Ting Lin, Cheng-Han Chiang and Hung-yi Lee · 2024
Later among the works it cites.
“Paralinguistics-enhanced large language modeling of spoken dialogue”
Guan-Ting Lin et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuan Gong et al · 2023
Cited alongside, same era.
“Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision”, 2023
Eugene Kharitonov et al · 2023
Cited alongside, same era.
“BigVGAN: A Universal Neural Vocoder with Large-Scale Training”, 2023
Sang-gil Lee et al · 2023
Cited alongside, same era.
“GPT-4 Technical Report” Accessed: 2025-07-11, https://openai.com/research/gpt-4 , 2023
OpenAI · 2023
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision”
Alec Radford et al · 2023
Cited alongside, same era.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2023
Cited alongside, same era.
“Audiopalm: A large language model that can speak and listen”
Paul Rubenstein et al · 2023
Cited alongside, same era.
Later among the works it cites.
“Spirit LM: Interleaved Spoken and Written Language Model”, 2024
Tu Nguyen et al · 2024
Later among the works it cites.
“Mmau: A massive multi-task audio understanding and reasoning benchmark”
S Sakshi et al · 2024
Later among the works it cites.
“Snac: Multi-scale neural audio codec”
Hubert Siuzdak, Florian Grötschla and Luca Lanzendörfer · 2024
Later among the works it cites.
“Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen llm”
Xiong Wang et al · 2024
Later among the works it cites.
“Maskgct: Zero-shot text-to-speech with masked generative codec transformer”
Yuancheng Wang et al · 2024
Later among the works it cites.
“Mini-omni: Language models can hear, talk while thinking in streaming”
Zhifei Xie and Changqiao Wu · 2024
Later among the works it cites.
“Mini-omni2: Towards open-source gpt-4o with vision, speech and duplex capabilities”
Zhifei Xie and Changqiao Wu · 2024
Later among the works it cites.
“Bigcodec: Pushing the limits of low-bitrate neural speech codec”
Detai Xin, Xu Tan, Shinnosuke Takamichi and Hiroshi Saruwatari · 2024
Later among the works it cites.
“Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot”
Aohan Zeng et al · 2024
Later among the works it cites.
“Minmo: A multimodal large language model for seamless voice interaction”
Qian Chen et al · 2025
Closest in time.
Ding Ding et al · 2025
Closest in time.
“LUCY: Linguistic Understanding and Control Yielding Early Stage of Her”, 2025
Heting Gao et al · 2025
Closest in time.
Sreyan Ghosh et al · 2025
Closest in time.
“Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models”, 2025
Arushi Goel et al · 2025
Closest in time.
“Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model”
Ailin Huang et al · 2025
Closest in time.
“Step-audio: Unified understanding and generation in intelligent speech interaction”
Ailin Huang et al · 2025
Closest in time.
“Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?”
Andrew Rouditchenko et al · 2025
Closest in time.
“Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens”
Xinsheng Wang et al · 2025
Closest in time.
“Qwen2.5-Omni Technical Report”, 2025
Jin Xu et al · 2025
Closest in time.
“URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models”, 2025
Ruiqi Yan et al · 2025
Closest in time.
“Minimax-speech: Intrinsic zero-shot text-to-speech with a learnable speaker encoder”
Bowen Zhang et al · 2025
Closest in time.
“Distinctive Feature Codec: Adaptive Segmentation for Efficient Speech Representation”
Xiangyu Zhang et al · 2025
Closest in time.