Fetching the paper…
Reading the bibliography…
Mainstream Audio Analytics models are trained to learn under the paradigm of one class label to many recordings focusing on one task.
“Machine hearing: An emerging field [exploratory dsp],”
Richard F Lyon, · 2010
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Earlier work this paper cites.
“Convolutional neural networks with binaural representations and background subtraction for acoustic scene classification,”
Yoonchang Han, Jeongsoo Park, and Kyogu Lee, · 2017
Earlier work this paper cites.
“Sound event detection in the dcase 2017 challenge,”
Annamaria Mesaros, Aleksandr Diment, Benjamin Elizalde, Toni Heittola, Emmanuel Vincent, Bhiksha Raj, and Tuomas Virtanen, · 2019
Earlier work this paper cites.
“AudioCaps: Generating Captions for Audios in The Wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Huggingface’s transformers: State-of-the-art natural language processing,”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al., · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“Sound event detection of weakly labelled data with cnn-transformer and automatic threshold optimization,”
Qiuqiang Kong, Yong Xu, Wenwu Wang, and Mark D Plumbley, · 2020
Cited alongside, same era.
“Clotho: an audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Cited alongside, same era.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley, · 2020
Cited alongside, same era.
Yu Zhang, Daniel S Park, Wei Han, James Qin, Anmol Gulati, Joel Shor, Aren Jansen, Yuanzhong Xu, Yanping Huang, Shibo Wang, et al., · 2021
Cited alongside, same era.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2021
“Broadcasted residual learning for efficient keyword spotting,”
Byeonggeun Kim, Simyung Chang, Jinkyu Lee, and Dooyong Sung, · 2021
Later among the works it cites.
“What is the ground truth? reliability of multi-annotator data for audio tagging,”
Irene Martín-Morató and Annamaria Mesaros, · 2021
Later among the works it cites.
“Lit: Zero-shot transfer with locked-image text tuning,”
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer, · 2021
Later among the works it cites.
“Wav2clip: Learning robust audio representations from clip,”
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello, · 2022
Closest in time.
“Audioclip: Extending clip to image, text and audio,”
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
“Florence: A new foundation model for computer vision,”
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al., · 2021
Cited alongside, same era.
“Scaling up visual and vision-language representation learning with noisy text supervision,”
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig, · 2021
Cited alongside, same era.
“Efficient training of audio transformers with patchout,”
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer, · 2021
Cited alongside, same era.
Joseph Turian, Jordie Shier, Humair Raj Khan, Bhiksha Raj, Björn W Schuller, Christian J Steinmetz, Colin Malloy, George Tzanetakis, Gissel Velarde, Kirk McNally, et al., · 2022
Closest in time.
“A proposal for multimodal emotion recognition using aural transformers and action units on ravdess dataset,”
Cristina Luna-Jiménez, Ricardo Kleinlein, David Griol, Zoraida Callejas, Juan M. Montero, and Fernando Fernández-Martínez, · 2022
Closest in time.
“Vocalsound: A dataset for improving human vocal sounds recognition,”
Yuan Gong, Jin Yu, and James Glass, · 2022
Closest in time.
“Fsd50k: An open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2022
Closest in time.