Fetching the paper…
Reading the bibliography…
While stochastic video prediction models enable future prediction under uncertainty, they mostly fail to model the complex dynamics of real-world scenes.
G. Johansson, “Visual perception of biological motion and a model for its analysis,” Perception & Psychophysics , vol. 14, no. 2, pp. 201–211, jun 1973
1973
Earlier work this paper cites.
C. Schüldt, I. Laptev, and B. Caputo, “Recognizing human actions: A local svm approach,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2004
2004
Earlier work this paper cites.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2012
2012
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The KITTI dataset,” International Journal of Robotics Research (IJRR) , 2013
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in Proc. of the International Conf. on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,” in Advances in Neural Information Processing Systems (NeurIPS) , 2014
2014
Earlier work this paper cites.
J. Walker, A. Gupta, and M. Hebert, “Dense optical flow prediction from a static image,” in Proc. of the IEEE International Conf. on Computer Vision (ICCV) , 2015
2015
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Zisserman, and k. kavukcuoglu, “Spatial transformer networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. of the International Conf. on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
X. Jia, B. De Brabandere, T. Tuytelaars, and L. V. Gool, “Dynamic filter networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2016
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
R. Garg, V. K. Bg, G. Carneiro, and I. Reid, “Unsupervised CNN for single view depth estimation: Geometry to the rescue,” in Proc. of the European Conf. on Computer Vision (ECCV) , 2016
2016
Earlier work this paper cites.
J. Y. Jason, A. W. Harley, and K. G. Derpanis, “Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness,” in Proc. of the European Conf. on Computer Vision (ECCV) Workshops , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
Z. Liu, R. A. Yeh, X. Tang, Y. Liu, and A. Agarwala, “Video frame synthesis using deep voxel flow,” in Proc. of the IEEE International Conf. on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
C. Lu, M. Hirsch, and B. Scholkopf, “Flexible spatio-temporal networks for video prediction,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , July 2017
2017
Earlier work this paper cites.
W. Lotter, G. Kreiman, and D. Cox, “Deep predictive coding networks for video prediction and unsupervised learning,” in Proc. of the International Conf. on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
C. Vondrick and A. Torralba, “Generating the future with adversarial transformers,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, “Unsupervised learning of depth and ego-motion from video,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
R. Villegas, J. Yang, Y. Zou, S. Sohn, X. Lin, and H. Lee, “Learning to generate long-term future via hierarchical prediction,” in international conference on machine learning . PMLR, 2017, pp. 3560–3569
2017
Earlier work this paper cites.
Z. Ren, J. Yan, B. Ni, B. Liu, X. Yang, and H. Zha, “Unsupervised deep learning for optical flow estimation,” in Proc. of the Conf. on Artificial Intelligence (AAAI) , 2017
2017
Earlier work this paper cites.
F. Ebert, C. Finn, A. X. Lee, and S. Levine, “Self-supervised visual planning with temporal skip connections,” in 1st Annual Conference on Robot Learning, CoRL 2017, Mountain View, California, USA, November 13-15, 2017, Proceedings , 2017
2017
Cited alongside, same era.
E. Denton and R. Fergus, “Stochastic video generation with a learned prior,” in Proc. of the International Conf. on Machine learning (ICML) , 2018
2018
Cited alongside, same era.
H. Zhan, R. Garg, C. Saroj Weerasekera, K. Li, H. Agarwal, and I. Reid, “Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
C. Wang, J. Miguel Buenaposada, R. Zhu, and S. Lucey, “Learning depth from monocular videos using direct methods,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
A. Ranjan, V. Jampani, L. Balles, K. Kim, D. Sun, J. Wulff, and M. J. Black, “Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Later among the works it cites.
C. Luo, Z. Yang, P. Wang, Y. Wang, W. Xu, R. Nevatia, and A. Yuille, “Every pixel counts++: Joint learning of geometry and motion with 3d holistic understanding,” IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI) , vol. 42, no. 10, 2019
2019
Later among the works it cites.
R. Villegas, A. Pathak, H. Kannan, D. Erhan, Q. V. Le, and H. Lee, “High fidelity video prediction with large stochastic recurrent neural networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2019
2019
Later among the works it cites.
M. Minderer, C. Sun, R. Villegas, F. Cole, K. P. Murphy, and H. Lee, “Unsupervised learning of object structure and dynamics from videos,” in Advances in Neural Information Processing Systems (NeurIPS) , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Mahjourian, M. Wicke, and A. Angelova, “Unsupervised learning of depth and ego-motion from monocular video using 3d geometric constraints,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
Z. Yin and J. Shi, “GeoNet: Unsupervised learning of dense depth, optical flow and camera pose,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
Y. Zou, Z. Luo, and J.-B. Huang, “DF-Net: Unsupervised joint learning of depth and flow using cross-task consistency,” in Proc. of the European Conf. on Computer Vision (ECCV) , 2018
2018
Cited alongside, same era.
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine, “Stochastic variational video prediction,” in Proc. of the International Conf. on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine, “Stochastic adversarial video prediction,” arXiv.org , 2018
2018
Cited alongside, same era.
S. Meister, J. Hur, and S. Roth, “Unflow: Unsupervised learning of optical flow with a bidirectional census loss,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2018
2018
Cited alongside, same era.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Cited alongside, same era.
F. A. Reda, G. Liu, K. J. Shih, R. Kirby, J. Barker, D. Tarjan, A. Tao, and B. Catanzaro, “Sdc-net: Video prediction using spatially-displaced convolution,” in Proc. of the European Conf. on Computer Vision (ECCV) , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
V. Casser, S. Pirk, R. Mahjourian, and A. Angelova, “Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos,” in Proc. of the Conf. on Artificial Intelligence (AAAI) , 2019
2019
Later among the works it cites.
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Towards accurate generative models of video: A new metric & challenges,” arXiv.org , 2019
2019
Later among the works it cites.
Y. Zhu, K. Sapra, F. A. Reda, K. J. Shih, S. Newsam, A. Tao, and B. Catanzaro, “Improving semantic segmentation via video propagation and label relaxation,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Later among the works it cites.
V. Guizilini, R. Ambrus, S. Pillai, A. Raventos, and A. Gaidon, “3D packing for self-supervised monocular depth estimation,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Later among the works it cites.
J. Gao, C. Sun, H. Zhao, Y. Shen, D. Anguelov, C. Li, and C. Schmid, “Vectornet: Encoding hd maps and agent dynamics from vectorized representation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 11 525–11 533
2020
Later among the works it cites.
J.-Y. Franceschi, E. Delasalles, M. Chen, S. Lamprier, and P. Gallinari, “Stochastic latent residual video prediction,” in Proc. of the International Conf. on Machine learning (ICML) , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
A. K. Akan, E. Erdem, A. Erdem, and F. Guney, “Slamp: Stochastic latent appearance and motion prediction,” in Proc. of the IEEE International Conf. on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
S. Safadoust and F. Güney, “Self-supervised monocular scene decomposition and depth estimation,” in International Conference on 3D Vision (3DV) , 2021, pp. 627–636
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Gu, C. Sun, and H. Zhao, “Densetnt: End-to-end trajectory prediction from dense goal sets,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 15 303–15 312
2021
Later among the works it cites.
A. Hu, Z. Murez, N. Mohan, S. Dudas, J. Hawke, V. Badrinarayanan, R. Cipolla, and A. Kendall, “FIERY: Future instance segmentation in bird’s-eye view from surround monocular cameras,” in Proceedings of the International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
O. Rybkin, K. Daniilidis, and S. Levine, “Simple and effective vae training with calibrated decoders,” in Proc. of the International Conf. on Machine learning (ICML) , 2021
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.