Fetching the paper…
Reading the bibliography…
Contrastive learning has shown remarkable success in the field of multimodal representation learning.
“Gradient-based learning applied to document recognition,”
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner, · 1998
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, and Richard Socher et al., · 2009
Earlier work this paper cites.
“Freesound technical demo,”
Frederic Font Corbera, Gerard Roma Trepat, and Xavier Serra, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello, · 2014
Earlier work this paper cites.
“ESC: dataset for environmental sound classification,”
Karol J. Piczak, · 2015
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke, Daniel P. W. Ellis, and Dylan Freedman et al., · 2017
Earlier work this paper cites.
“Accurate, large minibatch sgd: Training imagenet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2017
Earlier work this paper cites.
“Deep learning using rectified linear units (relu),”
Abien Fred Agarap, · 2018
Earlier work this paper cites.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized BERT pretraining approach,”
Yinhan Liu, Myle Ott, and Naman Goyal et al., · 2019
Earlier work this paper cites.
“Audiocaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Cited alongside, same era.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, and Turab Iqbal et al., · 2020
Cited alongside, same era.
“Clotho: an audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Cited alongside, same era.
“Exploring the limits of transfer learning with a unified text-to-text transformer,”
Colin Raffel, Noam Shazeer, and Adam Roberts et al., · 2020
Cited alongside, same era.
“Vggsound: A large-scale audio-visual dataset,”
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman, · 2020
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
“Attention bottlenecks for multimodal fusion,”
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun, · 2021
Later among the works it cites.
“Audioclip: Extending clip to image, text and audio,”
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel, · 2022
Closest in time.
“CLAP: learning audio concepts from natural language supervision,”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, · 2022
Closest in time.
“Audio retrieval with wavtext5k and CLAP training,”
Soham Deshmukh, Benjamin Elizalde, and Huaming Wang, · 2022
Closest in time.
“Text-to-audio retrieval via large-scale contrastive learning,”
Yusong Wu, Tianyu Zhang, and Ke Chen, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Jong Wook Kim, and Chris Hallacy et al., · 2021
Cited alongside, same era.
“Audio retrieval with natural language queries,”
Andreea-Maria Oncescu, A. Sophia Koepke, and João F. Henriques et al., · 2021
Cited alongside, same era.
“Audio retrieval with natural language queries: A benchmark study,”
A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata, and Samuel Albanie, · 2021
Cited alongside, same era.
“Swin transformer: Hierarchical vision transformer using shifted windows,”
Ze Liu, Yutong Lin, and Yue Cao et al., · 2021
Cited alongside, same era.
“Attentional feature fusion,”
Yimian Dai, Fabian Gieseke, and Stefan Oehmcke et al., · 2021
Cited alongside, same era.
Xinhao Mei, Xubo Liu, and Jianyuan Sun et al., · 2022
Closest in time.
“Wav2clip: Learning robust audio representations from clip,”
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello, · 2022
Closest in time.
“HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,”
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2022
Closest in time.
“Efficient training of audio transformers with patchout,”
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer, · 2022
Closest in time.
“FSD50K: an open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2022
Closest in time.