Fetching the paper…
Reading the bibliography…
Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy Lillicrap, Jonathan Hunt, Alexander Pritzel, Nicolas Heess, et al · 2015
Earlier work this paper cites.
ESC: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort Gemmeke, Daniel Ellis, Dylan Freedman, Aren Jansen, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Speech commands: A dataset for limited-vocabulary speech recognition
Pete Warden · 2018
Earlier work this paper cites.
Fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Earlier work this paper cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, et al · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, et al · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
PANNs: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark Plumbley · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision Transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, et al · 2021
Cited alongside, same era.
HTS-AT: A hierarchical token-semantic audio Transformer for sound classification and detection
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2022
Later among the works it cites.
WavLM: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Later among the works it cites.
BEATs: Audio pre-training with acoustic tokenizers
Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, and Furu Wei · 2022
Later among the works it cites.
SSAST: Self-supervised audio spectrogram Transformer
Yuan Gong, Cheng Lai, Yu-An Chung, and James Glass · 2022
Later among the works it cites.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jorn Hees, and Andreas Dengel · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Cited alongside, same era.
An empirical study of training self-supervised vision Transformers
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Cited alongside, same era.
AST: Audio spectrogram Transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Cited alongside, same era.
PSLA: Improving audio tagging with pretraining, sampling, labeling, and aggregation
Yuan Gong, Yu-An Chung, and James Glass · 2021
Cited alongside, same era.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Efficient training of audio Transformers with patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-Zadeh, and Gerhard Widmer · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows, 2021
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick · 2022
Later among the works it cites.
Masked autoencoders that listen
Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, et al · 2022
Later among the works it cites.
ATST: Audio representation learning with teacher-student Transformer
Xian Li and Xiaofei Li · 2022
Later among the works it cites.
Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2022
Later among the works it cites.
Conformer-based self-supervised learning for non-speech audio tasks
Sangeeta Srivastava, Yun Wang, Andros Tjandra, Anurag Kumar, Chunxi Liu, Kritika Singh, and Yatharth Saraf · 2022
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello · 2022
Later among the works it cites.
Efficient self-supervised learning with contextualized target representations for vision, speech and language
Alexei Baevski, Arun Babu, Wei-Ning Hsu, and Michael Auli · 2023
Later among the works it cites.
Masked spectrogram prediction for self-supervised audio pre-training
Dading Chong, Helin Wang, Peilin Zhou, and Qingcheng Zeng · 2023
Later among the works it cites.
Self-supervised audio teacher-student Transformer for both clip-level and frame-level tasks
Xian Li, Nian Shao, and Xiaofei Li · 2023
Later among the works it cites.
MT4SSL: Boosting self-supervised speech representation learning by integrating multiple targets
Ziyang Ma, Zhisheng Zheng, Changli Tang, Yujin Wang, and Xie Chen · 2023
Later among the works it cites.
Masked modeling duo: Learning representations by encouraging both networks to model the input
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2023
Later among the works it cites.