Fetching the paper…
Reading the bibliography…
This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space.
Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects
Rajesh PN Rao and Dana H Ballard · 1999
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
Karol J. Piczak · 2015
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krähenbühl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron Weiss, and Kevin Wilson · 2017
Earlier work this paper cites.
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
P. Warden · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Fast image caption generation with position alignment
Zheng-cong Fei · 2019
Earlier work this paper cites.
On the power of curriculum learning in training deep networks
Guy Hacohen and Daphna Weinshall · 2019
Earlier work this paper cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Earlier work this paper cites.
Competence-based curriculum learning for neural machine translation
Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabas Poczos, and Tom M Mitchell · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Earlier work this paper cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Earlier work this paper cites.
Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Iterative back modification for faster image captioning
Zhengcong Fei · 2020
Cited alongside, same era.
Learning representations by predicting bags of visual words
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, and Matthieu Cord · 2020
Cited alongside, same era.
Data-efficient image recognition with contrastive predictive coding
Olivier Henaff · 2020
Cited alongside, same era.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley · 2020
Cited alongside, same era.
Mae-ast: Masked autoencoding audio spectrogram transformer
Alan Baade, Puyuan Peng, and David Harwath · 2022
Later among the works it cites.
data2vec: A general framework for self-supervised learning in speech, vision and language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli · 2022
Later among the works it cites.
Vicregl: Self-supervised learning of local visual features
Adrien Bardes, Jean Ponce, and Yann LeCun · 2022
Later among the works it cites.
Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2022
Later among the works it cites.
Intra-instance vicreg: Bag of self-supervised image patch embedding explains the performance
Yubei Chen, Adrien Bardes, ZENGYI LI, and Yann LeCun · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Andy T. Liu, Shu-Wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee · 2020
Cited alongside, same era.
Voxceleb: Large-scale speaker verification in the wild
Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2021
Cited alongside, same era.
High fidelity visualization of what your self-supervised representation knows about
Florian Bordes, Randall Balestriero, and Pascal Vincent · 2021
Cited alongside, same era.
Partially non-autoregressive image captioning
Zhengcong Fei · 2021
Cited alongside, same era.
AST: audio spectrogram transformer
Yuan Gong, Yu-An Chung, and James R. Glass · 2021
Cited alongside, same era.
Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation
Yuan Gong, YuAn Chung, and James Glass · 2021
Cited alongside, same era.
Later among the works it cites.
Masked spectrogram prediction for self-supervised audio pre-training, 2022
Dading Chong, Helin Wang, Peilin Zhou, and Qingcheng Zeng · 2022
Later among the works it cites.
Deecap: Dynamic early exiting for efficient image captioning
Zhengcong Fei, Xu Yan, Shuhui Wang, and Qi Tian · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Yanghao Li, Kaiming He, et al · 2022
Later among the works it cites.
Cmkd: Cnn/transformer-based cross-model knowledge distillation for audio classification, 2022
Yuan Gong, Sameer Khurana, Andrew Rouditchenko, and James Glass · 2022
Later among the works it cites.
Ssast: Self-supervised audio spectrogram transformer
Yuan Gong, Cheng-I Lai, Yu-An Chung, and James Glass · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Masked autoencoders that listen
Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer · 2022
Later among the works it cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun · 2022
Later among the works it cites.
Juncheng B. Li, Shuhui Qu, Po-Yao Huang, and Florian Metze · 2022
Later among the works it cites.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2022
Later among the works it cites.
Improving multimodal speech recognition by data augmentation and speech representations
Dan Oneata and Horia Cucu · 2022
Later among the works it cites.
Robust self-supervised audio-visual speech recognition
Bowen Shi, Wei-Ning Hsu, and Abdelrahman Mohamed · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
ibot: Image bert pre-training with online tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong · 2022
Later among the works it cites.
Self-supervised learning from images with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas · 2023
Closest in time.
Efficient self-supervised learning with contextualized target representations for vision, speech and language
Alexei Baevski, Arun Babu, Wei-Ning Hsu, and Michael Auli · 2023
Closest in time.
Incorporating unlikely negative cues for distinctive image captioning
Zhengcong Fei and Junshi Huang · 2023
Closest in time.
Masked auto-encoders meet generative adversarial networks and beyond
Zhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang, Xiaoming Wei, and Xiaolin Wei · 2023
Closest in time.
Uncertainty-aware image captioning
Zhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang, Xiaoming Wei, and Xiaolin Wei · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.