Fetching the paper…
Reading the bibliography…
Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2019 · 1910
Earlier work this paper cites.
The Eighth Text REtrieval Conference (TREC-8) . National Institute of Standards and Technology (NIST)
E. Voorhees and D. Harman, editors. 1999 · 1999
Earlier work this paper cites.
Fsd50k: An open dataset of human-labeled sound events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. 2022 · 2010
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J. Piczak. 2015 · 2015
Earlier work this paper cites.
Query variations and their effect on comparing information retrieval systems
Guido Zuccon, Joao Palotti, and Allan Hanbury. 2016 · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
AudioCaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019 · 2019
Earlier work this paper cites.
The benefit of temporally-strong labels in audio event classification
Shawn Hershey, Daniel P W Ellis, Eduardo Fonseca, Aren Jansen, Caroline Liu, R Channing Moore, and Manoj Plakal. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
Generating datasets with pretrained language models
Timo Schick and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
Inpars: Data augmentation for information retrieval using large language models
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Cited alongside, same era.
Promptagator: Few-shot dense retrieval from 8 examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang. 2022 · 2022
Cited alongside, same era.
Audio retrieval with natural language queries: A benchmark study
A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata, and Samuel Albanie. 2023 · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu*, Ke Chen*, Tianyu Zhang*, Yuchen Hui*, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023 · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Leveraging llms for unsupervised dense retriever ranking
Ekaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, and Guido Zuccon. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Proceedings of the 7th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2022)
Mathieu Lagrange, Annamaria Mesaros, Thomas Pellegrini, Romain Serizel Gaël Richard, and Dan Stowell. 2022 · 2022
Cited alongside, same era.
On metric learning for audio-text cross-modal retrieval
Xinhao Mei, Xubo Liu, Jianyuan Sun, Mark D. Plumbley, and Wenwu Wang. 2022 · 2022
Cited alongside, same era.
On the evaluation metrics for paraphrase generation
Lingfeng Shen, Lemao Liu, Haiyun Jiang, and Shuming Shi. 2022 · 2022
Cited alongside, same era.
Clap learning audio concepts from natural language supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. 2023a · 2023
Cited alongside, same era.
Inpars-v2: Large language models as efficient dataset generators for information retrieval
Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira. 2023 · 2023
Cited alongside, same era.
Natural language supervision for general-purpose audio representations
Benjamin Elizalde, Soham Deshmukh, and Huaming Wang. 2023b
Cited in the paper.
Recap: Retrieval-augmented audio captioning
Sreyan Ghosh, Sonal Kumar, Chandra Kiran Reddy Evuru, Ramani Duraiswami, and Dinesh Manocha. 2024a
Cited in the paper.
Compa: Addressing the gap in compositional reasoning in audio-language models
Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Reddy Evuru, Ramaneswaran S, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha. 2024b
Cited in the paper.
Hyunjae Kim, Seunghyun Yoon, Trung Bui, Handong Zhao, Quan Tran, Franck Dernoncourt, and Jaewoo Kang. 2024 · 2024
Closest in time.
Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024 · 2024
Closest in time.
Gecko: Versatile text embeddings distilled from large language models
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Blair Chen, Daniel Cer, Jeremy R. Cole, Kai Hui, Michael Boratko, Rajvi Kapadia, Wen Ding, Yi Luan, Sai Meher Karthik Duddu, Gustavo Hernandez Abrego, Weiqiang Shi, Nithi Gupta, Aditya Kusupati, Prateek Jain, Siddhartha Reddy Jonnalagadda, Ming-Wei Chang, and Iftekhar Naim. 2024 · 2024
Closest in time.