Fetching the paper…
Reading the bibliography…
We present XKD, a novel self-supervised framework to learn meaningful representations from unlabelled videos.
On the convergence of adam and beyond
Reddi, S. J.; Kale, S.; and Kumar, S. 2019 · 1904
Earlier work this paper cites.
Learning spatiotemporal features via video and text pair discrimination
Li, T.; and Wang, L. 2020 · 2001
Earlier work this paper cites.
Audiovisual slowfast networks for video recognition
Xiao, F.; Lee, Y. J.; Grauman, K.; Malik, J.; and Feichtenhofer, C. 2020 · 2001
Earlier work this paper cites.
A kernel method for the two-sample-problem
Gretton, A.; Borgwardt, K.; Rasch, M.; Schölkopf, B.; and Smola, A. 2006 · 2006
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Kuehne, H.; Jhuang, H.; Garrote, E.; Poggio, T.; and Serre, T. 2011 · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
Piczak, K. J. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Abu-El-Haija, S.; Kothari, N.; Lee, J.; Natsev, P.; Toderici, G.; Varadarajan, B.; and Vijayanarasimhan, S. 2016 · 2016
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Aytar, Y.; Vondrick, C.; and Torralba, A. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016 · 2016
Earlier work this paper cites.
Look, listen and learn
Arandjelovic, R.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Tarvainen, A.; and Valpola, H. 2017 · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017 · 2017
Earlier work this paper cites.
Emotion recognition in speech using cross-modal transfer in the wild
Albanie, S.; Nagrani, A.; Vedaldi, A.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Self-supervised spatiotemporal feature learning via video rotation prediction
Jing, L.; Yang, X.; Liu, J.; and Tian, Y. 2018 · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Korbar, B.; Tran, D.; and Torresani, L. 2018 · 2018
Earlier work this paper cites.
Mixed Precision Training
Micikevicius, P.; Narang, S.; Alben, J.; Diamos, G.; Elsen, E.; Garcia, D.; Ginsburg, B.; Houston, M.; Kuchaiev, O.; Venkatesh, G.; et al. 2018 · 2018
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S.; Han, D.; Oh, S. J.; Chun, S.; Choe, J.; and Yoo, Y. 2019 · 2019
Cited alongside, same era.
Asr is all you need: Cross-modal distillation for lip reading
Afouras, T.; Chung, J. S.; and Zisserman, A. 2020 · 2020
Cited alongside, same era.
Self-Supervised MultiModal Versatile Networks
Alayrac, J.-B.; Recasens, A.; Schneider, R.; Arandjelovic, R.; Ramapuram, J.; De Fauw, J.; Smaira, L.; Dieleman, S.; and Zisserman, A. 2020 · 2020
Cited alongside, same era.
Self-Supervised Learning by Cross-Modal Audio-Video Clustering
Alwassel, H.; Mahajan, D.; Korbar, B.; Torresani, L.; Ghanem, B.; and Tran, D. 2020 · 2020
Cited alongside, same era.
Efficient training of audio transformers with patchout
Koutini, K.; Schlüter, J.; Eghbal-zadeh, H.; and Widmer, G. 2021 · 2021
Later among the works it cites.
Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
Min, S.; Dai, Q.; Xie, H.; Gan, C.; Zhang, Y.; and Wang, J. 2021 · 2021
Later among the works it cites.
Robust Audio-Visual Instance Discrimination
Morgado, P.; Misra, I.; and Vasconcelos, N. 2021 · 2021
Later among the works it cites.
Audio-visual instance discrimination with cross-modal agreement
Morgado, P.; Vasconcelos, N.; and Misra, I. 2021 · 2021
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Qian, R.; Meng, T.; Gong, B.; Yang, M.-H.; Wang, H.; Belongie, S.; and Cui, Y. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asano, Y. M.; Patrick, M.; Rupprecht, C.; and Vedaldi, A. 2020 · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020 · 2020
Cited alongside, same era.
Bootstrap Your Own Latent: A new approach to self-supervised learning
Grill, J.-B.; Strub, F.; Altché, F.; Tallec, C.; Richemond, P.; Buchatskaya, E.; Doersch, C.; Pires, B.; Guo, Z.; Azar, M.; et al. 2020 · 2020
Cited alongside, same era.
Self-supervised Co-training for Video Representation Learning
Han, T.; Xie, W.; and Zisserman, A. 2020 · 2020
Cited alongside, same era.
Active Contrastive Learning of Audio-Visual Video Representations
Ma, S.; Zeng, Z.; McDuff, D.; and Song, Y. 2020 · 2020
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
Miech, A.; Alayrac, J.-B.; Smaira, L.; Laptev, I.; Sivic, J.; and Zisserman, A. 2020 · 2020
Cited alongside, same era.
Broaden your views for self-supervised video learning
Recasens, A.; Luc, P.; Alayrac, J.-B.; Wang, L.; Strub, F.; Tallec, C.; Malinowski, M.; Pătrăucean, V.; Altché, F.; Valko, M.; et al. 2021 · 2021
Later among the works it cites.
Learning from the master: Distilling cross-modal advanced knowledge for lip reading
Ren, S.; Du, Y.; Lv, J.; Han, G.; and He, S. 2021 · 2021
Later among the works it cites.
Verma, P.; and Berger, J. 2021 · 2021
Later among the works it cites.
Bevt: Bert pretraining of video transformers
Wang, R.; Chen, D.; Wu, Z.; Chen, Y.; Dai, X.; Liu, M.; Jiang, Y.-G.; Zhou, L.; and Yuan, L. 2021 · 2021
Later among the works it cites.
Mae-ast: Masked autoencoding audio spectrogram transformer
Baade, A.; Peng, P.; and Harwath, D. 2022 · 2022
Closest in time.
MultiMAE: Multi-modal Multi-task Masked Autoencoders
Bachmann, R.; Mizrahi, D.; Atanov, A.; and Zamir, A. 2022 · 2022
Closest in time.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A.; Hsu, W.-N.; Xu, Q.; Babu, A.; Gu, J.; and Auli, M. 2022 · 2022
Closest in time.
Masked Spectrogram Prediction For Self-Supervised Audio Pre-Training
Chong, D.; Wang, H.; Zhou, P.; and Zeng, Q. 2022 · 2022
Closest in time.
Masked Autoencoders As Spatiotemporal Learners
Feichtenhofer, C.; Fan, H.; Li, Y.; and He, K. 2022 · 2022
Closest in time.
FSD50K: an open dataset of human-labeled sound events
Fonseca, E.; Favory, X.; Pons, J.; Font, F.; and Serra, X. 2022 · 2022
Closest in time.
Omnivore: A single model for many visual modalities
Girdhar, R.; Singh, M.; Ravi, N.; van der Maaten, L.; Joulin, A.; and Misra, I. 2022 · 2022
Closest in time.
Masked autoencoders that listen
Huang, P.-Y.; Xu, H.; Li, J.; Baevski, A.; Auli, M.; Galuba, W.; Metze, F.; and Feichtenhofer, C. 2022 · 2022
Closest in time.
Self-supervised learning for videos: A survey
Schiappa, M. C.; Rawat, Y. S.; and Shah, M. 2022 · 2022
Closest in time.
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Tong, Z.; Song, Y.; Wang, J.; and Wang, L. 2022 · 2022
Closest in time.
MaCLR: Motion-Aware Contrastive Learning of Representations for Videos
Xiao, F.; Tighe, J.; and Modolo, D. 2022 · 2022
Closest in time.
Scaling vision transformers to 22 billion parameters
Dehghani, M.; Djolonga, J.; Mustafa, B.; Padlewski, P.; Heek, J.; Gilmer, J.; Steiner, A. P.; Caron, M.; Geirhos, R.; Alabdulmohsin, I.; et al. 2023 · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K. V.; Joulin, A.; and Misra, I. 2023 · 2023
Closest in time.
Self-supervised audio-visual representation learning with relaxed cross-modal synchronicity
Sarkar, P.; and Etemad, A. 2023 · 2023
Closest in time.