Fetching the paper…
Reading the bibliography…
Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text.
Computational auditory scene analysis
G. J. Brown and M. Cooke · 1994
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
In two minds: dual-process accounts of reasoning
J. S. Evans · 2003
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Freesound technical demo
F. Font, G. Roma, and X. Serra · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Earlier work this paper cites.
Freesound datasets: a platform for the creation of open audio datasets
E. Fonseca, J. Pons Puig, X. Favory, F. Font Corbera, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra · 2017
Earlier work this paper cites.
Audio Set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Understanding, Explanation, and Scientific Knowledge
K. Khalifa · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2017
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
AudioCaps: Generating Captions for Audios in The Wild
C. D. Kim, B. Kim, H. Lee, and G. Kim · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, and Others · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman · 2020
Earlier work this paper cites.
Clotho: an Audio Captioning Dataset
K. Drossos, S. Lipping, and T. Virtanen · 2020
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, et al · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Winogrande: an adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Cited alongside, same era.
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
M. Abdin, J. Aneja, H. Awadalla, A. Awadallah, A. A. Awan, N. Bach, A. Bahree, A. Bakhtiari, J. Bao, H. Behl, et al · 2024
Later among the works it cites.
Paligemma: A versatile 3b vlm for transfer
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Cited alongside, same era.
HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Audio retrieval with natural language queries: A benchmark study
A. S. Koepke, A.-M. Oncescu, J. F. Henriques, Z. Akata, and S. Albanie · 2022
Cited alongside, same era.
Clotho-aqa: A crowdsourced dataset for audio question answering
S. Lipping, P. Sudarsanam, K. Drossos, and T. Virtanen · 2022
Cited alongside, same era.
Automated audio captioning: An overview of recent progress and new challenges
X. Mei, X. Liu, M. D. Plumbley, and W. Wang · 2022
Cited alongside, same era.
Diverse audio captioning via adversarial training
X. Mei, X. Liu, J. Sun, M. D. Plumbley, and W. Wang · 2022
Cited alongside, same era.
Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin, C. Zhou, and J. Zhou · 2024
Later among the works it cites.
Pam: Prompting audio-language models for audio quality assessment
S. Deshmukh, D. Alharthi, B. Elizalde, H. Gamper, M. Al Ismail, R. Singh, B. Raj, and H. Wang · 2024
Later among the works it cites.
Training audio captioning models without audio
S. Deshmukh, B. Elizalde, D. Emmanouilidou, B. Raj, R. Singh, and H. Wang · 2024
Later among the works it cites.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Later among the works it cites.
Natural language supervision for general-purpose audio representations
B. Elizalde, S. Deshmukh, and H. Wang · 2024
Later among the works it cites.
Listen, think, and understand
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass · 2024
Later among the works it cites.
Mobilellm: Optimizing sub-billion parameter language models for on-device use cases
Z. Liu, C. Zhao, F. Iandola, C. Lai, Y. Tian, I. Fedorov, Y. Xiong, E. Chang, Y. Shi, R. Krishnamoorthi, et al · 2024
Later among the works it cites.
Openelm: An efficient language model family with open-source training and inference framework
S. Mehta, M. H. Sekhavat, Q. Cao, M. Horton, Y. Jin, C. Sun, I. Mirzadeh, M. Najibi, D. Belenko, P. Zatloukal, et al · 2024
Later among the works it cites.
Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang · 2024
Later among the works it cites.
Mmau: A massive multi-task audio understanding and reasoning benchmark
S. Sakshi, U. Tyagi, S. Kumar, A. Seth, R. Selvakumar, O. Nieto, R. Duraiswami, S. Ghosh, and D. Manocha · 2024
Later among the works it cites.
Paligemma 2: A family of versatile vlms for transfer
A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y. Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, et al · 2024
Later among the works it cites.
SALMONN: Towards generic hearing abilities for large language models
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang · 2024
Later among the works it cites.
Computer audition: From task-specific machine learning to foundation models, 2024
A. Triantafyllopoulos, I. Tsangko, A. Gebhard, A. Mesaros, T. Virtanen, and B. Schuller · 2024
Later among the works it cites.
Smollm2: When smol goes big–data-centric training of a small language model
L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blázquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíček, A. P. Lajarín, V. Srivastav, et al · 2025
Closest in time.
ADIFF: Explaining audio difference using natural language
S. Deshmukh, S. Han, R. Singh, and B. Raj · 2025
Closest in time.
M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa, et al · 2025
Closest in time.
Improving Weakly Supervised Sound Event Detection with Self-Supervised Auxiliary Tasks
S. Deshmukh, B. Raj, and R. Singh · 2079
Closest in time.