Fetching the paper…
Reading the bibliography…
Over the past several years, the synchronization between audio and visual signals has been leveraged to learn richer audio-visual representations.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Unsupervised representation learning by predicting image rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Speech commands: A dataset for limited-vocabulary speech recognition
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer
Sachin Mehta and Mohammad Rastegari · 2021
Later among the works it cites.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2021
Later among the works it cites.
Contrastive audio-visual masked autoencoder
Yuan Gong, Andrew Rouditchenko, Alexander H Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R Glass · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Yanghao Li, Kaiming He, et al · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pete Warden · 2018
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Cvnets: High performance library for computer vision
Sachin Mehta, Farzad Abdolhosseini, and Mohammad Rastegari · 2022
Later among the works it cites.
Dianwen Ng, Yunqi Chen, Biao Tian, Qiang Fu, and Eng Siong Chng · 2022
Later among the works it cites.
Videomae v2: Scaling video masked autoencoders with dual masking
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao · 2023
Closest in time.
Diffusion models as masked autoencoders
Chen Wei, Karttikeya Mangalam, Po-Yao Huang, Yanghao Li, Haoqi Fan, Hu Xu, Huiyu Wang, Cihang Xie, Alan Yuille, and Christoph Feichtenhofer · 2023
Closest in time.
On the efficacy of multi-scale data samplers for vision applications
Elvis Nunez, Thomas Merth, Anish Prabhu, Mehrdad Farajtabar, Mohammad Rastegari, Sachin Mehta, and Maxwell Horton · 2023
Closest in time.