Fetching the paper…
Reading the bibliography…
Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database.
“Large-scale content-based audio retrieval from text queries,”
Gal Chechik, Eugene Ie, Martin Rehn, Samy Bengio, and Dick Lyon, · 2008
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Acoustic event search with an onomatopoeic query: measuring distance between onomatopoeic words and sounds,”
Shota Ikawa and Kunio Kashino, · 2018
Earlier work this paper cites.
“AudioCaps: Generating Captions for Audios in The Wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Cross modal audio search and retrieval with joint embeddings based on text and audio,”
Benjamin Elizalde, Shuayb Zarar, and Bhiksha Raj, · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Earlier work this paper cites.
“Clotho: an audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Earlier work this paper cites.
Never-Ending Learning of Sounds
Benjamin Martinez Elizalde, · 2020
Earlier work this paper cites.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley, · 2020
Cited alongside, same era.
“Multi-task learning for interpretable weakly labelled sound event detection,”
Soham Deshmukh, Bhiksha Raj, and Rita Singh, · 2020
Cited alongside, same era.
“What is the ground truth? reliability of multi-annotator data for audio tagging,”
Irene Martín-Morató and Annamaria Mesaros, · 2021
Cited alongside, same era.
“Audio retrieval with natural language queries,”
Andreea-Maria Oncescu, A Koepke, Joao F Henriques, Zeynep Akata, and Samuel Albanie, · 2021
Cited alongside, same era.
“AST: Audio Spectrogram Transformer,”
Yuan Gong, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Audio retrieval with natural language queries: A benchmark study,”
A. Sophia Koepke, Andreea-Maria Oncescu, Joao Henriques, Zeynep Akata, and Samuel Albanie, · 2022
Closest in time.
“Dcase 2022 challenge task 6b: Language-based audio retrieval,”
Huang Xie, Samuel Lipping, and Tuomas Virtanen, · 2022
Closest in time.
“On metric learning for audio-text cross-modal retrieval,”
Xinhao Mei, Xubo Liu, Jianyuan Sun, Mark D Plumbley, and Wenwu Wang, · 2022
Closest in time.
“Clap: Learning audio concepts from natural language supervision,”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, · 2022
Closest in time.
“Wav2clip: Learning robust audio representations from clip,”
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuan Gong, Cheng-I Jeff Lai, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Improving Weakly Supervised Sound Event Detection with Self-Supervised Auxiliary Tasks,”
Soham Deshmukh, Bhiksha Raj, and Rita Singh, · 2021
Cited alongside, same era.
Closest in time.
“Audioclip: Extending clip to image, text and audio,”
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel, · 2022
Closest in time.
“Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection,”
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2022
Closest in time.