Fetching the paper…
Reading the bibliography…
The quality of the image representations obtained from self-supervised learning depends strongly on the type of data augmentations used in the learning formulation.
Object recognition from local scale-invariant features
David G. Lowe · 1999
Earlier work this paper cites.
Recent advances in the automatic recognition of audiovisual speech
Gerasimos Potamianos, Chalapathy Neti, Guillaume Gravier, Ashutosh Garg, and Andrew W Senior · 2003
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Recurrence quantification analysis features for environmental sound recognition
Guido Roma, Waldo Nogueira, and Perfecto Herrera · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Discriminative unsupervised feature learning with exemplar convolutional neural networks
Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Environmental sound classification with convolutional neural networks
Karol J. Piczak · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J. Piczak · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Detection and classification of acoustic scenes and events
Dan Stowell, Dimitrios Giannoulis, Emmanouil Benetos, Mathieu Lagrange, and Mark D. Plumbley · 2015
Earlier work this paper cites.
Detection and classification of acoustic scenes and events
D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
Bolei Zhou, Aditya Khosla, Àgata Lapedriza, Aude Oliva, and Antonio Torralba · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Actionvlad: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan C. Russell · 2017
Earlier work this paper cites.
Learning spatio-temporal features with 3d residual networks for action recognition
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Understanding the effective receptive field in deep convolutional neural networks, 2017
Wenjie Luo, Yujia Li, Raquel Urtasun, and Richard Zemel · 2017
Earlier work this paper cites.
Learning from video and text via large-scale discriminative clustering
Antoine Miech, Jean-Baptiste Alayrac, Piotr Bojanowski, Ivan Laptev, and Josef Sivic · 2017
Earlier work this paper cites.
Unsupervised filterbank learning using convolutional restricted boltzmann machine for environmental sound classification
Hardik B. Sailor, Dharmesh M Agrawal, and Hemant A Patil · 2017
Earlier work this paper cites.
Dataset augmentation in feature space
V Terrance and W Taylor Graham · 2017
Earlier work this paper cites.
Long-term Temporal Convolutions for Action Recognition
Gül Varol, Ivan Laptev, and Cordelia Schmid · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Split-brain autoencoders: Unsupervised learning by cross-channel prediction
Richard Zhang, Phillip Isola, and Alexei A Efros · 2017
Earlier work this paper cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2018
Earlier work this paper cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Deep lip reading: A comparison of models and an online application
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Earlier work this paper cites.
Improving spatiotemporal self-supervision by deep reinforcement learning
Uta Buchler, Biagio Brattoli, and Bjorn Ommer · 2018
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Unsupervised representation learning by predicting image rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis · 2018
Cited alongside, same era.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adria Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Kenton Lee Jacob Devlin, Ming-Wei Chang and Kristina Toutanova · 2018
Cited alongside, same era.
Self-supervised spatiotemporal feature learning by video geometric transformations
Longlong Jing and Yingli Tian · 2018
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Later among the works it cites.
Soundspaces: Audio-visual navigation in 3d environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman · 2020
Later among the works it cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeff Wu, Heewoo Jun, Prafulla Dhariwal, David Luan, and Ilya Sutskever · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Later among the works it cites.
Self-supervised spatio-temporal representation learning using variable playback speed prediction
Hyeon Cho, Taehoon Kim, Hyung Jin Chang, and Wonjun Hwang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Learning and using the arrow of time
Donglai Wei, Joseph J Lim, Andrew Zisserman, and William T Freeman · 2018
Cited alongside, same era.
Randaugment: Practical data augmentation with no separate search
E. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le · 2020
Later among the works it cites.
Autoaugment: Learning augmentation policies from data
Ekin Dogus Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le · 2020
Later among the works it cites.
Virtex: Learning visual representations from textual annotations, 2020
Karan Desai and Justin Johnson · 2020
Later among the works it cites.
Crosstransformers: spatially-aware few-shot transfer
Carl Doersch, Ankush Gupta, and Andrew Zisserman · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Later among the works it cites.
Oops! predicting unintentional action in video
D. Epstein, Boyuan Chen, and Carl Vondrick · 2020
Later among the works it cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Later among the works it cites.
Listen to look: Action recognition by previewing audio
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko · 2020
Later among the works it cites.
Memory-augmented dense predictive coding for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Later among the works it cites.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Learning video representations by transforming time
S. Jenni, Givi Meishvili, and P. Favaro · 2020
Later among the works it cites.
Hard negative mixing for contrastive learning
Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus · 2020
Later among the works it cites.
Spatially aware multimodal transformers for textvqa
Yash Kant, Dhruv Batra, Peter Anderson, Alex Schwing, Devi Parikh, Jiasen Lu, and Harsh Agrawal · 2020
Later among the works it cites.
Video understanding as machine translation, 2020
Bruno Korbar, Fabio Petroni, Rohit Girdhar, and Lorenzo Torresani · 2020
Later among the works it cites.
Featmatch: Feature-based augmentation for semi-supervised learning
Chia-Wen Kuo, Chih-Yao Ma, Jia-Bin Huang, and Zsolt Kira · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
Center-wise local image mixture for contrastive representation learning, 2020
Hao Li, Xiaopeng Zhang, Ruoyu Sun, Hongkai Xiong, and Qi Tian · 2020
Later among the works it cites.
Prototypical contrastive learning of unsupervised representations
Junnan Li, Pan Zhou, Caiming Xiong, Richard Socher, and Steven CH Hoi · 2020
Later among the works it cites.
Learning spatiotemporal features via video and text pair discrimination
Tianhao Li and Limin Wang · 2020
Later among the works it cites.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
Learning audio-visual representations with active contrastive coding, 2020
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Later among the works it cites.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2020
Later among the works it cites.
Speech2action: Cross-modal supervision for action recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, and Andrew Zisserman · 2020
Later among the works it cites.
Multi-modal self-supervision from generalized data transformations
Mandela Patrick, Yuki Markus Asano, Ruth Fong, João F. Henriques, G. Zweig, and A. Vedaldi · 2020
Later among the works it cites.
Support-set bottlenecks for video-text representation learning, 2020
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, João Henriques, and Andrea Vedaldi · 2020
Later among the works it cites.
Evolving losses for unsupervised video representation learning
AJ Piergiovanni, Anelia Angelova, and Michael S. Ryoo · 2020
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, H. Wang, Serge J. Belongie, and Yin Cui · 2020
Later among the works it cites.
Learning visual representations with caption annotations
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus · 2020
Later among the works it cites.
Learning video representations from textual web supervision
Jonathan C. Stroud, D. Ross, Chen Sun, Jun Deng, R. Sukthankar, and C. Schmid · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Later among the works it cites.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
Hao Tan and Mohit Bansal · 2020
Later among the works it cites.
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Later among the works it cites.
What makes for good views for contrastive learning
Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola · 2020
Later among the works it cites.
Self-supervised video representation learning by pace prediction
Jiangliu Wang, Jianbo Jiao, and Y. Liu · 2020
Later among the works it cites.
Visual transformers: Token-based image representation and processing for computer vision, 2020
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Masayoshi Tomizuka, Kurt Keutzer, and Peter Vajda · 2020
Later among the works it cites.
Audiovisual slowfast networks for video recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?, 2020
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
Linchao Zhu and Yi Yang · 2020
Later among the works it cites.
Is space-time attention all you need for video understanding?, 2021
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Parameter efficient multimodal transformers for video representation learning
Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, and Yale Song · 2021
Closest in time.
Contrastive self-supervised learning of global-local audio-visual representations, 2021
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song · 2021
Closest in time.
Video transformer network, 2021
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.