Fetching the paper…
Reading the bibliography…
We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music.
Switchboard: Telephone speech corpus for research and development
J. J. Godfrey, E. C. Holliman, and J. McDaniel · 1992
Earlier work this paper cites.
The fisher corpus: A resource for the next generations of speech-to-text
C. Cieri, D. Miller, and K. Walker · 2004
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
P. Koehn · 2005
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan · 2008
Earlier work this paper cites.
Evaluation of algorithms using games: the case of music annotation
E. Law, K. West, M. Mandel, M. Bay, and J. Downie · 2010
Earlier work this paper cites.
The million song dataset
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere · 2011
Earlier work this paper cites.
The million song dataset
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere · 2011
Earlier work this paper cites.
Ted-lium: an automatic speech recognition dedicated corpus
A. Rousseau, P. Deléglise, and Y. Esteve · 2012
Earlier work this paper cites.
Medleydb: A multitrack dataset for annotation-intensive mir research
R. M. Bittner, J. Salamon, M. Tierney, M. Mauch, C. Cannam, and J. P. Bello · 2014
Earlier work this paper cites.
Chime-home: A dataset for sound source recognition in a domestic environment
P. Foster, S. Sigtia, S. Krstulovic, J. Barker, and M. D. Plumbley · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Earlier work this paper cites.
Fma: A dataset for music analysis
M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson · 2016
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan · 2017
Earlier work this paper cites.
Freesound datasets: A platform for the creation of open audio datasets
E. Fonseca, J. Pons, X. Favory, F. Font, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
The musdb18 corpus for music separation
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner · 2017
Earlier work this paper cites.
The omg-emotion behavior dataset
P. Barros, N. Churamani, E. Lakomkin, H. Siqueira, A. Sutherland, and S. Wermter · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
J. S. Chung, A. Nagrani, and A. Zisserman · 2018
Earlier work this paper cites.
Toward truly personal chatbots: on the development of custom conversational assistants
F. Daniel, M. Matera, V. Zaccaria, and A. Dell’Orto · 2018
Earlier work this paper cites.
Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve · 2018
Earlier work this paper cites.
An open source emotional speech corpus for human robot interaction applications
J. James, L. Tian, and C. I. Watson · 2018
Earlier work this paper cites.
A multi-device dataset for urban acoustic scene classification
A. Mesaros, T. Heittola, and T. Virtanen · 2018
Earlier work this paper cites.
Meld: A multimodal multi-party dataset for emotion recognition in conversations
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea · 2018
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
C. D. Kim, B. Kim, H. Lee, and G. Kim · 2019
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber · 2020
Earlier work this paper cites.
Sonyc-ust-v2: An urban sound tagging dataset with spatiotemporal context
M. Cartwright, J. Cramer, A. E. M. Mendez, Y. Wang, H.-H. Wu, V. Lostanlen, M. Fuentes, G. Dove, C. Mydlarz, J. Salamon, et al · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman · 2020
Earlier work this paper cites.
100,000 podcasts: A spoken English document corpus
A. Clifton, S. Reddy, Y. Yu, A. Pappu, R. Rezapour, H. Bonab, M. Eskevich, G. Jones, J. Karlgren, B. Carterette, and R. Jones · 2020
Earlier work this paper cites.
Clotho: An audio captioning dataset
K. Drossos, S. Lipping, and T. Virtanen · 2020
Earlier work this paper cites.
The msp-conversation corpus
L. Martinez-Lucas, M. Abdelwahab, and C. Busso · 2020
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert · 2020
Earlier work this paper cites.
Music4all: A new music database and its applications
I. A. P. Santana, F. Pinhelli, J. Donini, L. Catharin, R. B. Mangolin, V. D. Feltrim, M. A. Domingues, et al · 2020
Earlier work this paper cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
G. Chen, S. Chai, G. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, et al · 2021
Earlier work this paper cites.
Fsd50k: an open dataset of human-labeled sound events
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra · 2021
Earlier work this paper cites.
Audioclip: Extending clip to image, text and audio, 2021
A. Guzhov, F. Raue, J. Hees, and A. Dengel · 2021
Earlier work this paper cites.
The benefit of temporally-strong labels in audio event classification
S. Hershey, D. P. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal · 2021
Earlier work this paper cites.
Diversity and bias in audio captioning datasets
I. M. Morato and A. Mesaros · 2021
Earlier work this paper cites.
Audio retrieval with natural language queries
A.-M. Oncescu, A. Koepke, J. F. Henriques, Z. Akata, and S. Albanie · 2021
Earlier work this paper cites.
P. K. O’Neill, V. Lavrukhin, S. Majumdar, V. Noroozi, Y. Zhang, O. Kuchaiev, J. Balam, Y. Dovzhenko, K. Freyberg, M. D. Shulman, et al · 2021
Earlier work this paper cites.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux · 2021
Cited alongside, same era.
Wav2clip: Learning robust audio representations from clip, 2021
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello · 2021
Cited alongside, same era.
Beats: Audio pre-training with acoustic tokenizers, 2022
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, and F. Wei · 2022
Cited alongside, same era.
Audio retrieval with wavtext5k and clap training
S. Deshmukh, B. Elizalde, and H. Wang · 2022
Cited alongside, same era.
Clap: Learning audio concepts from natural language supervision, 2022
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang · 2022
Cited alongside, same era.
Voicebench: Benchmarking llm-based voice assistants
Y. Chen, X. Yue, C. Zhang, X. Gao, R. T. Tan, and H. Li · 2024
Later among the works it cites.
Qwen2-audio technical report, 2024
Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin, C. Zhou, and J. Zhou · 2024
Later among the works it cites.
M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou · 2024
Later among the works it cites.
Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities, 2024
S. Ghosh, S. Kumar, A. Seth, C. K. R. Evuru, U. Tyagi, Sakshi, O. Nieto, R. Duraiswami, and D. Manocha · 2024
Later among the works it cites.
Audio dialogues: Dialogues dataset for audio and music understanding
A. Goel, Z. Kong, R. Valle, and B. Catanzaro · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mmdialog: A large-scale multi-turn dialogue dataset towards multi-modal open-domain conversation
J. Feng, Q. Sun, C. Xu, P. Zhao, Y. Yang, C. Tao, D. Zhao, and Q. Lin · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2022
Cited alongside, same era.
Cochlscene: Acquisition of acoustic scene data using crowdsourcing
I.-Y. Jeong and J. Park · 2022
Cited alongside, same era.
Audio retrieval with natural language queries: A benchmark study
A. S. Koepke, A.-M. Oncescu, J. F. Henriques, Z. Akata, and S. Albanie · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
D. Lee, C. Kim, S. Kim, M. Cho, and W.-S. Han · 2022
Cited alongside, same era.
Learning to answer questions in dynamic audio-visual scenarios
G. Li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu · 2022
Cited alongside, same era.
Clotho-aqa: A crowdsourced dataset for audio question answering
S. Lipping, P. Sudarsanam, K. Drossos, and T. Virtanen · 2022
Cited alongside, same era.
Later among the works it cites.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Later among the works it cites.
Video recap: Recursive captioning of hour-long videos
M. M. Islam, N. Ho, X. Yang, T. Nagarajan, L. Torresani, and G. Bertasius · 2024
Later among the works it cites.
Miradata: A large-scale video dataset with long durations and structured captions
X. Ju, Y. Gao, Z. Zhang, Z. Yuan, X. Wang, A. Zeng, Y. Xiong, Q. Xu, and Y. Shan · 2024
Later among the works it cites.
Libriheavy: A 50,000 hours asr corpus with punctuation casing and context
W. Kang, X. Yang, Z. Yao, F. Kuang, Y. Yang, L. Guo, L. Lin, and D. Povey · 2024
Later among the works it cites.
Efficient generative modeling with residual vector quantization-based tokens
J. Kim, T. Moon, K. Lee, and J. Cho · 2024
Later among the works it cites.
Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities, 2024
Z. Kong, A. Goel, R. Badlani, W. Ping, R. Valle, and B. Catanzaro · 2024
Later among the works it cites.
Nv-embed: Improved techniques for training llms as generalist embedding models
C. Lee, R. Roy, M. Xu, J. Raiman, M. Shoeybi, B. Catanzaro, and W. Ping · 2024
Later among the works it cites.
S. Leng, Y. Xing, Z. Cheng, Y. Zhou, H. Zhang, X. Li, D. Zhao, S. Lu, C. Miao, and L. Bing · 2024
Later among the works it cites.
Music understanding llama: Advancing text-to-music generation with question answering and captioning
S. Liu, A. S. Hussain, C. Sun, and Y. Shan · 2024
Later among the works it cites.
Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang · 2024
Later among the works it cites.
Position: Levels of agi for operationalizing progress on the path to agi
M. R. Morris, J. Sohl-Dickstein, N. Fiedel, T. Warkentin, A. Dafoe, A. Faust, C. Farabet, and S. Legg · 2024
Later among the works it cites.
Sonics: Synthetic or not–identifying counterfeit songs
M. A. Rahman, Z. I. A. Hakim, N. H. Sarker, B. Paul, and S. A. Fattah · 2024
Later among the works it cites.
Mmau: A massive multi-task audio understanding and reasoning benchmark
S. Sakshi, U. Tyagi, S. Kumar, A. Seth, R. Selvakumar, O. Nieto, R. Duraiswami, S. Ghosh, and D. Manocha · 2024
Later among the works it cites.
Muchomusic: Evaluating music understanding in multimodal audio-language models
B. Weck, I. Manco, E. Benetos, E. Quinton, G. Fazekas, and D. Bogdanov · 2024
Later among the works it cites.
Mini-omni: Language models can hear, talk while thinking in streaming
Z. Xie and C. Wu · 2024
Later among the works it cites.
Llava-o1: Let vision language models reason step-by-step
G. Xu, P. Jin, L. Hao, Y. Song, L. Sun, and L. Yuan · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al · 2024
Later among the works it cites.
Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras
A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, J. Bao, A. Benhaim, M. Cai, V. Chaudhary, C. Chen, et al · 2025
Closest in time.
Audio entailment: Assessing deductive reasoning for audio understanding
S. Deshmukh, S. Han, H. Bukhari, B. Elizalde, H. Gamper, R. Singh, and B. Raj · 2025
Closest in time.
Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities, 2025
S. Ghosh, Z. Kong, S. Kumar, S. Sakshi, J. Kim, W. Ping, R. Valle, D. Manocha, and B. Catanzaro · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Step-audio: Unified understanding and generation in intelligent speech interaction
A. Huang, B. Wu, B. Wang, C. Yan, C. Hu, C. Feng, F. Tian, F. Shen, J. Li, M. Chen, et al · 2025
Closest in time.
DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors
K. Lee, D. W. Kim, J. Kim, S. Chung, and J. Cho · 2025
Closest in time.
Reinforcement learning outperforms supervised fine-tuning: A case study on audio question answering
G. Li, J. Liu, H. Dinkel, Y. Niu, J. Zhang, and J. Luan · 2025
Closest in time.
Baichuan-audio: A unified framework for end-to-end speech interaction
T. Li, J. Liu, T. Zhang, Y. Fang, D. Pan, M. Wang, Z. Liang, Z. Li, M. Lin, G. Dong, et al · 2025
Closest in time.
Audio-cot: Exploring chain-of-thought reasoning in large audio language model
Z. Ma, Z. Chen, Y. Wang, E. S. Chng, and X. Chen · 2025
Closest in time.
Mmar: A challenging benchmark for deep reasoning in speech, audio, music, and their mix, 2025
Z. Ma, Y. Ma, Y. Zhu, C. Yang, Y.-W. Chao, R. Xu, W. Chen, Y. Chen, Z. Chen, J. Cong, K. Li, K. Li, S. Li, X. Li, X. Li, Z. Lian, Y. Liang, M. Liu, Z. Niu, T. Wang, Y. Wang, Y. Wang, Y. Wu, G. Yang, J. Yu, R. Yuan, Z. Zheng, Z. Zhou, H. Zhu, W. Xue, E. Benetos, K. Yu, E.-S. Chng, and X. Chen · 2025
Closest in time.
Mqad: A large-scale question answering dataset for training music large language models
Z. Ouyang, J.-C. Wang, D. Zhang, B. Chen, S. Li, and Q. Lin · 2025
Closest in time.
Mmsu: A massive multi-task spoken language understanding and reasoning benchmark
D. Wang, J. Wu, J. Li, D. Yang, X. Chen, T. Zhang, and H. Meng · 2025
Closest in time.
Multimodal chain-of-thought reasoning: A comprehensive survey
Y. Wang, S. Wu, Y. Zhang, S. Yan, Z. Liu, J. Luo, and H. Fei · 2025
Closest in time.
Audio-reasoner: Improving reasoning capability in large audio language models
Z. Xie, M. Lin, Z. Liu, P. Wu, S. Yan, and C. Miao · 2025
Closest in time.
Qwen2. 5-omni technical report
J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, et al · 2025
Closest in time.
Sound-vecaps: Improving audio generation with visually enhanced captions
Y. Yuan, D. Jia, X. Zhuang, Y. Chen, Z. Chen, Y. Wang, Y. Wang, X. Liu, X. Kang, M. D. Plumbley, et al · 2025
Closest in time.
Are you really listening? boosting perceptual awareness in music-qa benchmarks
Y. Zang, S. O’Brien, T. Berg-Kirkpatrick, J. McAuley, and Z. Novack · 2025
Closest in time.