Fetching the paper…
Reading the bibliography…
Contrastive learning has nearly closed the gap between supervised and self-supervised learning of image representations, and has also been explored for videos.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Accelerating t-sne using tree-based algorithms
Laurens Van Der Maaten · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C. Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Earlier work this paper cites.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Improving spatiotemporal self-supervision by deep reinforcement learning
Uta Buchler, Biagio Brattoli, and Bjorn Ommer · 2018
Earlier work this paper cites.
Towards good practice for action recognition with spatiotemporal 3d convolutions
K. Hara, H. Kataoka, and Y. Satoh · 2018
Earlier work this paper cites.
Self-supervised spatiotemporal feature learning via video rotation prediction
Longlong Jing, Xiaodong Yang, Jingen Liu, and Yingli Tian · 2018
Earlier work this paper cites.
Resound: Towards action recognition without representation bias
Yingwei Li, Yi Li, and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Learning spatiotemporal 3d convolution with video order self-supervision
Tomoyuki Suzuki, Takahiro Itazuri, Kensho Hara, and Hirokatsu Kataoka · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Learning and using the arrow of time
Donglai Wei, Joseph J. Lim, Andrew Zisserman, and William T. Freeman · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
Video Jigsaw: Unsupervised learning of spatiotemporal context for video action recognition
Unaiza Ahsan, Rishi Madhok, and Irfan Essa · 2019
Earlier work this paper cites.
Learning representations by maximizing mutual information across views
Philip Bachman, R. Devon Hjelm, and William Buchwalter · 2019
Earlier work this paper cites.
Why can’t i dance in the mall? learning to mitigate scene bias in action recognition
Jinwoo Choi, Chen Gao, Joseph C.E. Messou, and Jia-Bin Huang · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Cited alongside, same era.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Cited alongside, same era.
Learning video representations using contrastive bidirectional transformer
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Cited alongside, same era.
Self-supervised video representation learning by pace prediction
Jiangliu Wang, Jianbo Jiao, and Yun-Hui Liu · 2020
Later among the works it cites.
Self-supervised video representation learning by maximizing mutual information
Fei Xue, Hongbing Ji, Wenbo Zhang, and Yi Cao · 2020
Later among the works it cites.
Video representation learning with visual tempo consistency
Ceyuan Yang, Yinghao Xu, Bo Dai, and Bolei Zhou · 2020
Later among the works it cites.
Unsupervised learning from video with deep neural embeddings
Chengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark, and Daniel Yamins · 2020
Later among the works it cites.
Unsupervised video representation learning by bidirectional feature prediction
Nadine Behrmann, Jurgen Gall, and Mehdi Noroozi · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Cited alongside, same era.
Self-supervised learning of audio-visual objects from video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Cited alongside, same era.
Self-Supervised Learning by Cross-Modal Audio-Video Clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Cited alongside, same era.
Can temporal information help with contrastive self-supervised learning?
Yutong Bai, Haoqi Fan, Ishan Misra, Ganesh Venkatesh, Yongyi Lu, Yuyin Zhou, Qihang Yu, Vikas Chandra, and Alan Yuille · 2020
Cited alongside, same era.
Speednet: Learning the speediness in videos
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel · 2020
Cited alongside, same era.
Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Rspnet: Relative speed perception for unsupervised video representation learning
Peihao Chen, Deng Huang, Dongliang He, Xiang Long, Runhao Zeng, Shilei Wen, Mingkui Tan, and Chuang Gan · 2021
Closest in time.
Self-supervised visual learning by variable playback speeds prediction of a video
Hyeon Cho, Taehoon Kim, Hyung Jin Chang, and Wonjun Hwang · 2021
Closest in time.
” knights”: First place submission for vipriors21 action recognition challenge at iccv 2021
Ishan Dave, Naman Biyani, Brandon Clark, Rohit Gupta, Yogesh Rawat, and Mubarak Shah · 2021
Closest in time.
A large-scale study on unsupervised spatiotemporal representation learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He · 2021
Closest in time.
Motion-augmented self-training for video recognition at smaller scale
Kirill Gavrilyuk, Mihir Jain, Ilia Karmanov, and Cees GM Snoek · 2021
Closest in time.
Self-supervised video representation learning with constrained spatiotemporal jigsaw, 2021
Yuqi Huo, Mingyu Ding, Haoyu Lu, Zhiwu Lu, Tao Xiang, Ji-Rong Wen, Ziyuan Huang, Jianwen Jiang, Shiwei Zhang, Mingqian Tang, Songfang Huang, and Ping Luo · 2021
Closest in time.
Time-equivariant contrastive video representation learning
Simon Jenni and Hailin Jin · 2021
Closest in time.
Temporally coherent embeddings for self-supervised video representation learning
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, and Peyman Moghadam · 2021
Closest in time.
Videomoco: Contrastive video representation learning with temporally adversarial examples
Tian Pan, Yibing Song, Tianyu Yang, Wenhao Jiang, and Wei Liu · 2021
Closest in time.
Multi-modal self-supervision from generalized data transformations, 2021
Mandela Patrick, Yuki Asano, Polina Kuznetsova, Ruth Fong, Joao F. Henriques, Geoffrey Zweig, and Andrea Vedaldi · 2021
Closest in time.
Enhancing self-supervised video representation learning via multi-level feature optimization
Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, and Weiyao Lin · 2021
Closest in time.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2021
Closest in time.
Self-supervised temporal learning, 2021
Hao Shao, Yu Liu, and Hongsheng Li · 2021
Closest in time.
Enhancing unsupervised video representation learning by decoupling the scene and the motion
Jinpeng Wang, Yuting Gao, Ke Li, Xinyang Jiang, Xiaowei Guo, Rongrong Ji, and Xing Sun · 2021
Closest in time.
Self-supervised video representation learning by uncovering spatio-temporal statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Wei Liu, and Yun-Hui Liu · 2021
Closest in time.
Seco: Exploring sequence supervision for unsupervised representation learning
Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, and Tao Mei · 2021
Closest in time.
Vipriors 2: Visual inductive priors for data-efficient deep learning challenges
Attila Lengyel, Robert-Jan Bruintjes, Marcos Baptista Rios, Osman Semih Kayhan, Davide Zambrano, Nergis Tomen, and Jan van Gemert · 2022
Closest in time.