Fetching the paper…
Reading the bibliography…
Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of LLMs.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
Reimers, N. and Gurevych, I · 2004
Earlier work this paper cites.
The gtzan dataset: Its contents, its faults, their effects on evaluation, and its future use
Sturm, B. L · 2013
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Cao, H., Cooper, D. G., Keutmann, M. K., Gur, R. C., Nenkova, A., and Verma, R · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Salamon, J., Jacoby, C., and Bello, J. P · 2014
Earlier work this paper cites.
Chime-home: A dataset for sound source recognition in a domestic environment
Foster, P., Sigtia, S., Krstulovic, S., Barker, J., and Plumbley, M. D · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Fma: A dataset for music analysis
Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X · 2016
Earlier work this paper cites.
Accelerating recurrent neural network training using sequence bucketing and multi-gpu data parallelization
Khomenko, V., Shyshkov, O., Radyvonenko, O., and Bokhan, K · 2016
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders, 2017
Engel, J., Resnick, C., Roberts, A., Dieleman, S., Eck, D., Simonyan, K., and Norouzi, M · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings
Lotfian, R. and Busso, C · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
The emotional voices database: Towards controlling the emotion dimension in voice generation systems
Adigwe, A., Tits, N., Haddad, K. E., Ostadabbas, S., and Dutoit, T · 2018
Earlier work this paper cites.
The omg-emotion behavior dataset
Barros, P., Churamani, N., Lakomkin, E., Siqueira, H., Sutherland, A., and Wermter, S · 2018
Earlier work this paper cites.
An open source emotional speech corpus for human robot interaction applications
James, J., Tian, L., and Watson, C · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Livingstone, S. R. and Russo, F. A · 2018
Earlier work this paper cites.
Meld: A multimodal multi-party dataset for emotion recognition in conversations
Poria, S., Hazarika, D., Majumder, N., Naik, G., Cambria, E., and Mihalcea, R · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Bogdanov, D., Won, M., Tovstogan, P., Porter, A., and Serra, X · 2019
Earlier work this paper cites.
Sonyc urban sound tagging (sonyc-ust): A multilabel dataset from an urban acoustic sensor network
Cartwright, M., Mendez, A. E. M., Cramer, A., Lostanlen, V., Dove, G., Wu, H.-H., Salamon, J., Nov, O., and Bello, J · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Johnson, J., Douze, M., and Jégou, H · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Medley-solos-DB: a cross-collection dataset for musical instrument recognition, February 2019
Lostanlen, V., Cella, C.-E., Bittner, R., and Essid, S · 2019
Earlier work this paper cites.
Musdb18-hq - an uncompressed version of musdb18, August 2019
Rafii, Z., Liutkus, A., Stöter, F.-R., Mimilakis, S. I., and Bittner, R · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Clotho: An audio captioning dataset
Drossos, K., Lipping, S., and Virtanen, T · 2020
Cited alongside, same era.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Cited alongside, same era.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Cited alongside, same era.
Toronto emotional speech set (TESS), 2020
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Musiclm: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Later among the works it cites.
Chen, Z., Huang, H., Andrusenko, A., Hrinchuk, O., Puvvada, K. C., Li, J., Ghosh, S., Balam, J., and Ginsburg, B · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pichora-Fuller, M. K. and Dupuis, K · 2020
Cited alongside, same era.
Fsd50k: an open dataset of human-labeled sound events
Fonseca, E., Favory, X., Pons, J., Font, F., and Serra, X · 2021
Cited alongside, same era.
Ast: Audio spectrogram transformer
Gong, Y., Chung, Y.-A., and Glass, J · 2021
Cited alongside, same era.
The benefit of temporally-strong labels in audio event classification
Hershey, S., Ellis, D. P., Fonseca, E., Jansen, A., Liu, C., Moore, R. C., and Plakal, M · 2021
Cited alongside, same era.
Diversity and bias in audio captioning datasets
Martin Morato, I. and Mesaros, A · 2021
Cited alongside, same era.
Audio retrieval with natural language queries
Oncescu, A.-M., Koepke, A., Henriques, J. F., Akata, Z., and Albanie, S · 2021
Cited alongside, same era.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, M., Menick, J. L., Cabi, S., Eslami, S., Vinyals, O., and Hill, F · 2021
Cited alongside, same era.
Chu, Y., Xu, J., Zhou, X., Yang, Q., Zhang, S., Yan, Z., Zhou, C., and Zhou, J · 2023
Later among the works it cites.
Pengi: An audio language model for audio tasks
Deshmukh, S., Elizalde, B., Singh, R., and Wang, H · 2023
Later among the works it cites.
Lp-musiccaps: Llm-based pseudo music captioning
Doh, S., Choi, K., Lee, J., and Nam, J · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Later among the works it cites.
Botchat: Evaluating llms’ capabilities of having multi-turn dialogues
Duan, H., Wei, J., Wang, C., Liu, H., Fang, Y., Zhang, S., Lin, D., and Chen, K · 2023
Later among the works it cites.
Clap learning audio concepts from natural language supervision
Elizalde, B., Deshmukh, S., Al Ismail, M., and Wang, H · 2023
Later among the works it cites.
Llark: A multimodal foundation model for music
Gardner, J., Durand, S., Stoller, D., and Bittner, R. M · 2023
Later among the works it cites.
Recap: Retrieval-augmented audio captioning
Ghosh, S., Kumar, S., Evuru, C. K. R., Duraiswami, R., and Manocha, D · 2023
Later among the works it cites.
Imagebind-llm: Multi-modality instruction tuning
Han, J., Zhang, R., Shao, W., Gao, P., Xu, P., Xiao, H., Zhang, K., Liu, C., Wen, S., Guo, Z., et al · 2023
Later among the works it cites.
An exploration of in-context learning for speech language model
Hsu, M.-H., Chang, K.-W., Li, S.-W., and Lee, H.-y · 2023
Later among the works it cites.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Huang, R., Li, M., Yang, D., Shi, J., Chang, X., Ye, Z., Wu, Y., Hong, Z., Huang, J., Liu, J., et al · 2023
Later among the works it cites.
Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration
Lyu, C., Wu, M., Wang, L., Huang, X., Liu, B., Du, Z., Shi, S., and Tu, Z · 2023
Later among the works it cites.
Mei, X., Meng, C., Liu, H., Kong, Q., Ko, T., Zhao, C., Plumbley, M. D., Zou, Y., and Wang, W · 2023
Later among the works it cites.
Anymal: An efficient and scalable any-modality augmented language model
Moon, S., Madotto, A., Lin, Z., Nagarajan, T., Smith, M., Jain, S., Yeh, C.-F., Murugesan, P., Heidari, P., Liu, Y., et al · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Later among the works it cites.
Nonspeech7k dataset: Classification and analysis of human non-speech sound
Rashid, M. M., Li, G., and Du, C · 2023
Later among the works it cites.
Audiopalm: A large language model that can speak and listen
Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., Bapna, A., Borsos, Z., Quitry, F. d. C., Chen, P., Badawy, D. E., Han, W., Kharitonov, E., et al · 2023
Later among the works it cites.
Zero-shot audio captioning with audio-language model guidance and audio context keywords
Salewski, L., Fauth, S., Koepke, A., and Akata, Z · 2023
Later among the works it cites.
Can whisper perform speech-based in-context learning
Wang, S., Yang, C.-H. H., Wu, J., and Zhang, C · 2023
Later among the works it cites.
A foundation model for music informatics
Won, M., Hung, Y.-N., and Le, D · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Later among the works it cites.
Re-vilm: Retrieval-augmented visual language model for zero and few-shot image captioning
Yang, Z., Ping, W., Liu, Z., Korthikanti, V., Nie, W., Huang, D.-A., Fan, L., Yu, Z., Lan, S., Li, B., et al · 2023
Later among the works it cites.
Chatbridge: Bridging modalities with large language model as a language catalyst
Zhao, Z., Guo, L., Yue, T., Chen, S., Shao, S., Zhu, X., Yuan, Z., and Liu, J · 2023
Later among the works it cites.