Fetching the paper…
Reading the bibliography…
Attempt to fully discover the temporal diversity and chronological characteristics for self-supervised video representation learning, this work takes advantage of the temporal dependencies within videos and further proposes a novel self-supervised method named Temporal Contrastive Graph Learning (TCGL).
Segregation of form, color, movement, and depth: anatomy, physiology, and perception
Margaret Livingstone and David Hubel · 1988
Earlier work this paper cites.
Neural mechanisms of form and motion processing in the primate visual system
David C Van Essen and Jack L Gallant · 1994
Earlier work this paper cites.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
The visual brain in action
David Milner and Mel Goodale · 2006
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
Alexander Klaser, Marcin Marszałek, and Cordelia Schmid · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Visual perception
Tom Cornsweet · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Dense trajectories and motion boundary descriptors for action recognition
Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Stap: Spatial-temporal attention-aware pooling for action recognition
Tam V Nguyen, Zheng Song, and Shuicheng Yan · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
Limin Wang, Yu Qiao, and Xiaoou Tang · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta · 2015
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros · 2016
Earlier work this paper cites.
Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice
Xiaojiang Peng, Limin Wang, Xingxing Wang, and Yu Qiao · 2016
Cited alongside, same era.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Cited alongside, same era.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Cited alongside, same era.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Adaptively connected neural networks
Guangrun Wang, Keze Wang, and Liang Lin · 2019
Later among the works it cites.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
Explainable video action reasoning via prior knowledge and state transitions
Tao Zhuo, Zhiyong Cheng, Peng Zhang, Yongkang Wong, and Mohan Kankanhalli · 2019
Later among the works it cites.
Speednet: Learning the speediness in videos
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Colorization as a proxy task for visual understanding
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2017
Cited alongside, same era.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Improving spatiotemporal self-supervision by deep reinforcement learning
Uta Buchler, Biagio Brattoli, and Bjorn Ommer · 2018
Cited alongside, same era.
Hierarchically learned view-invariant representations for cross-view action recognition
Yang Liu, Zhaoyang Lu, Jing Li, and Tao Yang · 2018
Cited alongside, same era.
Global temporal representation based cnns for infrared action recognition
Yang Liu, Zhaoyang Lu, Jing Li, Tao Yang, and Chao Yao · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al · 2020
Later among the works it cites.
Graphcl: Contrastive self-supervised learning of graph representations
Hakim Hafidi, Mounir Ghogho, Philippe Ciblat, and Ananthram Swami · 2020
Later among the works it cites.
The visual system
J Hans and Johannes RM Cruysberg · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Later among the works it cites.
Temporal contrastive pretraining for video action recognition
Guillaume Lorre, Jaonary Rabarisoa, Astrid Orcesi, Samia Ainouz, and Stephane Canu · 2020
Later among the works it cites.
Video cloze procedure for self-supervised spatio-temporal learning
Dezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang, Can Ma, Qixiang Ye, and Weiping Wang · 2020
Later among the works it cites.
Gcc: Graph contrastive coding for graph neural network pre-training
Jiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang, Hongxia Yang, Ming Ding, Kuansan Wang, and Jie Tang · 2020
Later among the works it cites.
Self-supervised video representation learning using inter-intra contrastive framework
Li Tao, Xueting Wang, and Toshihiko Yamasaki · 2020
Later among the works it cites.
Self-supervised video representation learning by pace prediction
Jiangliu Wang, Jianbo Jiao, and Yun-Hui Liu · 2020
Later among the works it cites.
A comprehensive survey on graph neural networks
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S Yu Philip · 2020
Later among the works it cites.
Explore video clip order with self-supervised and curriculum learning for video applications
Jun Xiao, Lin Li, Dejing Xu, Chengjiang Long, Jian Shao, Shifeng Zhang, Shiliang Pu, and Yueting Zhuang · 2020
Later among the works it cites.
Video playback rate perception for self-supervised spatio-temporal representation learning
Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou, and Qixiang Ye · 2020
Later among the works it cites.
Temporal reasoning graph for activity recognition
Jingran Zhang, Fumin Shen, Xing Xu, and Heng Tao Shen · 2020
Later among the works it cites.
Deep learning on graphs: A survey
Ziwei Zhang, Peng Cui, and Wenwu Zhu · 2020
Later among the works it cites.
Deep graph contrastive representation learning
Yanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu, Shu Wu, and Liang Wang · 2020
Later among the works it cites.
Self-supervised video representation learning by uncovering spatio-temporal statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Wei Liu, and Yun-Hui Liu · 2021
Closest in time.