Fetching the paper…
Reading the bibliography…
This paper addresses the problem of self-supervised video representation learning from a new perspective -- by video pace prediction.
Watamaniuk, S.N., Duchon, A.: The human visual system averages speed information. Vision research 32
1992
Earlier work this paper cites.
Giese, M.A., Poggio, T.: Neural mechanisms for the recognition of biological movements. Nature Reviews Neuroscience 4
2003
Earlier work this paper cites.
Laptev, I.: On space-time interest points. IJCV 64
2005
Earlier work this paper cites.
Dalal, N., Triggs, B., Schmid, C.: Human detection using oriented histograms of flow and appearance. In: ECCV (2006)
2006
Earlier work this paper cites.
Klaser, A., Marszałek, M., Schmid, C.: A spatio-temporal descriptor based on 3d-gradients. In: BMVC (2008)
2008
Earlier work this paper cites.
Laptev, I., Marszalek, M., Schmid, C., Rozenfeld, B.: Learning realistic human actions from movies. In: CVPR (2008)
2008
Earlier work this paper cites.
Gutmann, M., Hyvärinen, A.: Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In: AISTATS (2010)
2010
Earlier work this paper cites.
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: Hmdb: a large video database for human motion recognition. In: ICCV (2011)
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
Wang, H., Schmid, C.: Action recognition with improved trajectories. In: ICCV (2013)
2013
Earlier work this paper cites.
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: Large-scale video classification with convolutional neural networks. In: CVPR (2014)
2014
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: NeruIPS (2014)
2014
Earlier work this paper cites.
Caba Heilbron, F., Escorcia, V., Ghanem, B., Carlos Niebles, J.: Activitynet: A large-scale video benchmark for human activity understanding. In: CVPR (2015)
2015
Earlier work this paper cites.
Doersch, C., Gupta, A., Efros, A.A.: Unsupervised visual representation learning by context prediction. In: ICCV (2015)
2015
Earlier work this paper cites.
Srivastava, N., Mansimov, E., Salakhudinov, R.: Unsupervised learning of video representations using lstms. In: ICML (2015)
2015
Earlier work this paper cites.
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: ICCV (2015)
2015
Earlier work this paper cites.
Wang, X., Gupta, A.: Unsupervised learning of visual representations using videos. In: ICCV (2015)
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
2016
Earlier work this paper cites.
Misra, I., Zitnick, C.L., Hebert, M.: Shuffle and learn: unsupervised learning using temporal order verification. In: ECCV (2016)
2016
Earlier work this paper cites.
Noroozi, M., Favaro, P.: Unsupervised learning of visual representations by solving jigsaw puzzles. In: ECCV (2016)
2016
Earlier work this paper cites.
Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., Efros, A.A.: Context encoders: Feature learning by inpainting. In: CVPR (2016)
2016
Earlier work this paper cites.
Shou, Z., Wang, D., Chang, S.F.: Temporal action localization in untrimmed videos via multi-stage cnns. In: CVPR (2016)
2016
Earlier work this paper cites.
Vondrick, C., Pirsiavash, H., Torralba, A.: Generating videos with scene dynamics. In: NeurIPS (2016)
2016
Earlier work this paper cites.
Zhang, R., Isola, P., Efros, A.A.: Colorful image colorization. In: ECCV (2016)
2016
Cited alongside, same era.
Misra, I., Zitnick, C.L., Hebert, M.: Shuffle and learn: unsupervised learning usingtemporal order verification. In: ECCV. pp. 527–544. Springer (2016)
2016
Cited alongside, same era.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: CVPR (2017)
2017
Cited alongside, same era.
Doersch, C., Zisserman, A.: Multi-task self-supervised visual learning. In: ICCV (2017)
2017
Cited alongside, same era.
Fernando, B., Bilen, H., Gavves, E., Gould, S.: Self-supervised video representation learning with odd-one-out networks. In: CVPR (2017)
2017
Cited alongside, same era.
Wang, B., Ma, L., Zhang, W., Liu, W.: Reconstruction network for video captioning. In: CVPR (2018)
2018
Later among the works it cites.
Wang, J., Jiang, W., Ma, L., Liu, W., Xu, Y.: Bidirectional attentive fusion with context gating for dense video captioning. In: CVPR (2018)
2018
Later among the works it cites.
Wu, Z., Xiong, Y., Yu, S.X., Lin, D.: Unsupervised feature learning via non-parametric instance discrimination. In: CVPR (2018)
2018
Later among the works it cites.
Xie, S., Sun, C., Huang, J., Tu, Z., Murphy, K.: Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In: ECCV (2018)
2018
Later among the works it cites.
Jiang, H., Sun, D., Jampani, V., Yang, M.H., Learned-Miller, E., Kautz, J.: Superslomo: High quality estimation of multiple intermediate frames for video interpola-tion. In: CVPR. pp. 9000–9008 (2018)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Lee, H.Y., Huang, J.B., Singh, M., Yang, M.H.: Unsupervised representation learning by sorting sequences. In: ICCV (2017)
2017
Cited alongside, same era.
Shou, Z., Chan, J., Zareian, A., Miyazawa, K., Chang, S.F.: Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos. In: CVPR (2017)
2017
Cited alongside, same era.
Zagoruyko, S., Komodakis, N.: Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In: ICLR (2017)
2017
Cited alongside, same era.
Lee, H.Y., Huang, J.B., Singh, M., Yang, M.H.: Unsupervised representation learn-ing by sorting sequences. In: ICCV. pp. 667–676 (2017)
2017
Cited alongside, same era.
Zagoruyko, S., Komodakis, N.: Paying more attention to attention: Improving theperformance of convolutional neural networks via attention transfer. In: ICLR (2017)
2017
Cited alongside, same era.
Buchler, U., Brattoli, B., Ommer, B.: Improving spatiotemporal self-supervision by deep reinforcement learning. In: ECCV) (2018)
2018
Cited alongside, same era.
2018
Later among the works it cites.
Wei, D., Lim, J.J., Zisserman, A., Freeman, W.T.: Learning and using the arrow oftime. In: CVPR. pp. 8052–8060 (2018)
2018
Later among the works it cites.
2019
Later among the works it cites.
Bachman, P., Hjelm, R.D., Buchwalter, W.: Learning representations by maximizing mutual information across views. In: NeurIPS (2019)
2019
Later among the works it cites.
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition. In: ICCV (2019)
2019
Later among the works it cites.
Han, T., Xie, W., Zisserman, A.: Video representation learning by dense predictive coding. In: ICCV Workshops (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Hussein, N., Gavves, E., Smeulders, A.W.: Timeception for complex action recognition. In: CVPR (2019)
2019
Later among the works it cites.
Kim, D., Cho, D., Kweon, I.S.: Self-supervised video representation learning with space-time cubic puzzles. In: AAAI (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Wang, J., Jiao, J., Bao, L., He, S., Liu, Y., Liu, W.: Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics. In: CVPR (2019)
2019
Later among the works it cites.
Xu, D., Xiao, J., Zhao, Z., Shao, J., Xie, D., Zhuang, Y.: Self-supervised spatiotemporal learning via video clip order prediction. In: CVPR (2019)
2019
Later among the works it cites.
Benaim, S., Ephrat, A., Lang, O., Mosseri, I., Freeman, W.T., Rubinstein, M., Irani, M., Dekel, T.: Speednet: Learning the speediness in videos. In: CVPR (2020)
2020
Closest in time.
2020
Closest in time.
Han, T., Xie, W., Zisserman, A.: Memory-augmented dense predictive coding for video representation learning. In: ECCV (2020)
2020
Closest in time.
Jenni, S., Meishvili, G., Favaro, P.: Video representation learning by recognizing temporal transformations. In: ECCV (2020)
2020
Closest in time.
2020
Closest in time.
Yao, Y., Liu, C., Luo, D., Zhou, Y., Ye, Q.: Video playback rate perception for self-supervised spatio-temporal representation learning. In: CVPR (2020)
2020
Closest in time.