Fetching the paper…
Reading the bibliography…
Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments.
The acquisition of expert performance as problem solving: Construction and modification of mediating mechanisms through deliberate practice
Ericsson, K. A · 2003
Earlier work this paper cites.
Evaluation of algorithms using games: The case of music tagging
Law, E., West, K., Mandel, M. I., Bay, M., and Downie, J. S · 2009
Earlier work this paper cites.
Determinantal point processes for machine learning
Kulesza, A., Taskar, B., et al · 2012
Earlier work this paper cites.
Freesound technical demo
Font, F., Roma, G., and Serra, X · 2013
Earlier work this paper cites.
The gtzan dataset: Its contents, its faults, their effects on evaluation, and its future use
Sturm, B. L · 2013
Earlier work this paper cites.
Medleydb: A multitrack dataset for annotation-intensive mir research
Bittner, R. M., Salamon, J., Tierney, M., Mauch, M., Cannam, C., and Bello, J. P · 2014
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Cao, H., Cooper, D. G., Keutmann, M. K., Gur, R. C., Nenkova, A., and Verma, R · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Salamon, J., Jacoby, C., and Bello, J. P · 2014
Earlier work this paper cites.
Chime-home: A dataset for sound source recognition in a domestic environment
Foster, P., Sigtia, S., Krstulovic, S., Barker, J., and Plumbley, M. D · 2015
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
Piczak, K. J · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., and Vijayanarasimhan, S · 2016
Earlier work this paper cites.
Fma: A dataset for music analysis
Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings
Lotfian, R. and Busso, C · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
The emotional voices database: Towards controlling the emotion dimension in voice generation systems
Adigwe, A., Tits, N., Haddad, K. E., Ostadabbas, S., and Dutoit, T · 2018
Earlier work this paper cites.
The omg-emotion behavior dataset
Barros, P., Churamani, N., Lakomkin, E., Siqueira, H., Sutherland, A., and Wermter, S · 2018
Earlier work this paper cites.
An open source emotional speech corpus for human robot interaction applications
James, J., Tian, L., and Watson, C · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Livingstone, S. R. and Russo, F. A · 2018
Earlier work this paper cites.
Meld: A multimodal multi-party dataset for emotion recognition in conversations
Poria, S., Hazarika, D., Majumder, N., Naik, G., Cambria, E., and Mihalcea, R · 2018
Earlier work this paper cites.
Espnet: End-to-end speech processing toolkit
Watanabe, S., Hori, T., Karita, S., Hayashi, T., Nishitoba, J., Unno, Y., Soplin, N. E. Y., Heymann, J., Wiesner, M., Chen, N., et al · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Bogdanov, D., Won, M., Tovstogan, P., Porter, A., and Serra, X · 2019
Earlier work this paper cites.
Sonyc urban sound tagging (sonyc-ust): A multilabel dataset from an urban acoustic sensor network
Cartwright, M., Mendez, A. E. M., Cramer, A., Lostanlen, V., Dove, G., Wu, H.-H., Salamon, J., Nov, O., and Bello, J · 2019
Earlier work this paper cites.
A comparison of end-to-end models for long-form speech recognition
Chiu, C.-C., Han, W., Zhang, Y., Pang, R., Kishchenko, S., Nguyen, P., Narayanan, A., Liao, H., Zhang, S., Kannan, A., et al · 2019
Earlier work this paper cites.
Tau urban acoustic scenes 2019 openset, development dataset.”
Heittola, T., Mesaros, A., and Virtanen, T · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Medley-solos-DB: a cross-collection dataset for musical instrument recognition, February 2019
Lostanlen, V., Cella, C.-E., Bittner, R., and Essid, S · 2019
Earlier work this paper cites.
Musdb18-hq-an uncompressed version of musdb18
Rafii, Z., Liutkus, A., Stöter, F.-R., Mimilakis, S. I., and Bittner, R · 2019
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset, 2020
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Earlier work this paper cites.
Clotho: An audio captioning dataset
Drossos, K., Lipping, S., and Virtanen, T · 2020
Earlier work this paper cites.
Toronto emotional speech set (tess)
Pichora-Fuller, M. K. and Dupuis, K · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Fsd50k: an open dataset of human-labeled sound events
Fonseca, E., Favory, X., Pons, J., Font, F., and Serra, X · 2021
Earlier work this paper cites.
Ast: Audio spectrogram transformer
Gong, Y., Chung, Y.-A., and Glass, J · 2021
Cited alongside, same era.
The benefit of temporally-strong labels in audio event classification
Hershey, S., Ellis, D. P., Fonseca, E., Jansen, A., Liu, C., Moore, R. C., and Plakal, M · 2021
Cited alongside, same era.
Diversity and bias in audio captioning datasets
Martin Morato, I. and Mesaros, A · 2021
Cited alongside, same era.
Macs - multi-annotator captioned soundscapes, July 2021
Morato, I. M. and Mesaros, A · 2021
Cited alongside, same era.
Audio retrieval with natural language queries
Oncescu, A.-M., Koepke, A., Henriques, J. F., Akata, Z., and Albanie, S · 2021
Cited alongside, same era.
Audio-text models do not yet leverage natural language
Wu, H.-H., Nieto, O., Bello, J. P., and Salamon, J · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Abdin, M., Aneja, J., Behl, H., Bubeck, S., Eldan, R., Gunasekar, S., Harrison, M., Hewett, R. J., Javaheripi, M., Kauffmann, P., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Chen, K., Du, X., Zhu, B., Ma, Z., Berg-Kirkpatrick, T., and Dubnov, S · 2022
Cited alongside, same era.
Audio retrieval with wavtext5k and clap training
Deshmukh, S., Elizalde, B., and Wang, H · 2022
Cited alongside, same era.
Fsd50k: An open dataset of human-labeled sound events, 2022
Fonseca, E., Favory, X., Pons, J., Font, F., and Serra, X · 2022
Cited alongside, same era.
Audioclip: Extending clip to image, text and audio
Guzhov, A., Raue, F., Hees, J., and Dengel, A · 2022
Cited alongside, same era.
Towards reasoning in large language models: A survey
Huang, J. and Chang, K. C.-C · 2022
Cited alongside, same era.
Instruction pre-training: Language models are supervised multitask learners
Cheng, D., Gu, Y., Huang, S., Bi, J., Huang, M., and Wei, F · 2024
Later among the works it cites.
Chu, Y., Xu, J., Yang, Q., Wei, H., Wei, X., Guo, Z., Leng, Y., Lv, Y., He, J., Lin, J., et al · 2024
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Later among the works it cites.
Audio entailment: Assessing deductive reasoning for audio understanding
Deshmukh, S., Han, S., Bukhari, H., Elizalde, B., Gamper, H., Singh, R., and Raj, B · 2024
Later among the works it cites.
GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities
Ghosh, S., Kumar, S., Seth, A., Evuru, C. K. R., Tyagi, U., Sakshi, S., Nieto, O., Duraiswami, R., and Manocha, D · 2024
Later among the works it cites.
Listen, think, and understand
Gong, Y., Luo, H., Liu, A. H., Karlinsky, L., and Glass, J. R · 2024
Later among the works it cites.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Huang, R., Li, M., Yang, D., Shi, J., Chang, X., Ye, Z., Wu, Y., Hong, Z., Huang, J., Liu, J., et al · 2024
Later among the works it cites.
Video recap: Recursive captioning of hour-long videos
Islam, M. M., Ho, N., Yang, X., Nagarajan, T., Torresani, L., and Bertasius, G · 2024
Later among the works it cites.
Miradata: A large-scale video dataset with long durations and structured captions
Ju, X., Gao, Y., Zhang, Z., Yuan, Z., Wang, X., Zeng, A., Xiong, Y., Xu, Q., and Shan, Y · 2024
Later among the works it cites.
Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Kong, Z., Goel, A., Badlani, R., Ping, W., Valle, R., and Catanzaro, B · 2024
Later among the works it cites.
Sila: Signal-to-language augmentation for enhanced control in text-to-audio generation
Kumar, S., Seetharaman, P., Salamon, J., Manocha, D., and Nieto, O · 2024
Later among the works it cites.
Nv-embed: Improved techniques for training llms as generalist embedding models
Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W · 2024
Later among the works it cites.
MERT: Acoustic music understanding model with large-scale self-supervised training
LI, Y., Yuan, R., Zhang, G., Ma, Y., Chen, X., Yin, H., Xiao, C., Lin, C., Ragni, A., Benetos, E., Gyenge, N., Dannenberg, R., Liu, R., Chen, W., Xia, G., Shi, Y., Huang, W., Wang, Z., Guo, Y., and Fu, J · 2024
Later among the works it cites.
Music understanding llama: Advancing text-to-music generation with question answering and captioning
Liu, S., Hussain, A. S., Sun, C., and Shan, Y · 2024
Later among the works it cites.
Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
Mei, X., Meng, C., Liu, H., Kong, Q., Ko, T., Zhao, C., Plumbley, M. D., Zou, Y., and Wang, W · 2024
Later among the works it cites.
Position: Levels of AGI for operationalizing progress on the path to AGI
Morris, M. R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., and Legg, S · 2024
Later among the works it cites.
Sonics: Synthetic or not–identifying counterfeit songs
Rahman, M. A., Hakim, Z. I. A., Sarker, N. H., Paul, B., and Fattah, S. A · 2024
Later among the works it cites.
Mmau: A massive multi-task audio understanding and reasoning benchmark
Sakshi, S., Tyagi, U., Kumar, S., Seth, A., Selvakumar, R., Nieto, O., Duraiswami, R., Ghosh, S., and Manocha, D · 2024
Later among the works it cites.
Do audio-language models understand linguistic variations?
Selvakumar, R., Kumar, S., Giri, H. K., Anand, N., Seth, A., Ghosh, S., and Manocha, D · 2024
Later among the works it cites.
SALMONN: Towards generic hearing abilities for large language models
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., MA, Z., and Zhang, C · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Later among the works it cites.
Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video
Venkataramanan, S., Rizve, M. N., Carreira, J., Asano, Y. M., and Avrithis, Y · 2024
Later among the works it cites.
Muchomusic: Evaluating music understanding in multimodal audio-language models
Weck, B., Manco, I., Benetos, E., Quinton, E., Fazekas, G., and Bogdanov, D · 2024
Later among the works it cites.
Longvila: Scaling long-context visual language models for long videos
Xue, F., Chen, Y., Li, D., Hu, Q., Zhu, L., Li, X., Fang, Y., Tang, H., Yang, S., Liu, Z., et al · 2024
Later among the works it cites.
Sound-vecaps: Improving audio generation with visual enhanced captions
Yuan, Y., Jia, D., Zhuang, X., Chen, Y., Liu, Z., Chen, Z., Wang, Y., Wang, Y., Liu, X., Kang, X., et al · 2024
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2024
Later among the works it cites.
Reclap: Improving zero shot audio classification by describing sounds
Ghosh, S., Kumar, S., Evuru, C. K. R., Nieto, O., Duraiswami, R., and Manocha, D · 2025
Closest in time.
Longvlm: Efficient long video understanding via large language models
Weng, Y., Han, M., He, H., Chang, X., and Zhuang, B · 2025
Closest in time.