Fetching the paper…
Reading the bibliography…
The predictive learning of spatiotemporal sequences aims to generate future images by learning from the historical context, where the visual dynamics are believed to have modular structures that can be learned with compositional subsystems.
G. E. Hinton, “Distributed representations,” 1984
1984
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Cognitive modeling , vol. 5, no. 3, p. 1, 1988
1988
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proceedings of the IEEE , vol. 78, no. 10, pp. 1550–1560, 1990
1990
Earlier work this paper cites.
A. Krogh and J. Vedelsby, “Neural network ensembles, cross validation, and active learning,” in NeurIPS , 1995, pp. 231–238
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
S. Bengio and Y. Bengio, “Taking on the curse of dimensionality in joint distributions using neural networks,” IEEE Transactions on Neural Networks , vol. 11, no. 3, pp. 550–557, 2000
2000
Earlier work this paper cites.
C. Schuldt, I. Laptev, and B. Caputo, “Recognizing human actions: a local svm approach,” in ICPR , 2004
2004
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of machine learning research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
Z.-H. Zhou, “Ensemble learning.” Encyclopedia of biometrics , vol. 1, pp. 270–273, 2009
2009
Earlier work this paper cites.
I. Sutskever, J. Martens, and G. E. Hinton, “Generating text with recurrent neural networks,” in ICML , 2011
2011
Earlier work this paper cites.
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in ICML , 2014
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NeurIPS , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NeurIPS , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo, “Convolutional LSTM network: A machine learning approach for precipitation nowcasting,” in NeurIPS , 2015, pp. 802–810
2015
Earlier work this paper cites.
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond short snippets: Deep networks for video classification,” in CVPR , 2015, pp. 4694–4702
2015
Earlier work this paper cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015, pp. 2625–2634
2015
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in NeurIPS , 2015, pp. 1171–1179
2015
Earlier work this paper cites.
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh, “Action-conditional video prediction using deep networks in atari games,” in NeurIPS , 2015, pp. 2863–2871
2015
Earlier work this paper cites.
N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using LSTMs,” in ICML , 2015, pp. 843–852
2015
Earlier work this paper cites.
E. Denton, S. Chintala, R. Fergus et al. , “Deep generative image models using a Laplacian pyramid of adversarial networks,” in NeurIPS , 2015, pp. 1486–1494
2015
Earlier work this paper cites.
S. Sukhbaatar, J. Weston, R. Fergus et al. , “End-to-end memory networks,” in NeurIPS , 2015, pp. 2440–2448
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
C. Finn, I. Goodfellow, and S. Levine, “Unsupervised learning for physical interaction through video prediction,” in NeurIPS , 2016, pp. 64–72
2016
Earlier work this paper cites.
M. Mathieu, C. Couprie, and Y. LeCun, “Deep multi-scale video prediction beyond mean square error,” in ICLR , 2016
2016
Earlier work this paper cites.
T. Xue, J. Wu, K. Bouman, and B. Freeman, “Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks,” in NeurIPS , 2016
2016
Earlier work this paper cites.
V. Patraucean, A. Handa, and R. Cipolla, “Spatio-temporal video autoencoder with differentiable memory,” in ICLR Workshop , 2016
2016
Cited alongside, same era.
A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwińska, S. G. Colmenarejo, E. Grefenstette, T. Ramalho, J. Agapiou et al. , “Hybrid computing using a neural network with dynamic external memory,” Nature , vol. 538, no. 7626, pp. 471–476, 2016
2016
Cited alongside, same era.
B. De Brabandere, X. Jia, T. Tuytelaars, and L. Van Gool, “Dynamic filter networks,” in NeurIPS , 2016, pp. 667–675
2016
Cited alongside, same era.
Y. Wang, M. Long, J. Wang, Z. Gao, and S. Y. Philip, “PredRNN: Recurrent neural networks for predictive learning using spatiotemporal lstms,” in NeurIPS , 2017, pp. 879–888
2017
Cited alongside, same era.
J. Wu, E. Lu, P. Kohli, B. Freeman, and J. Tenenbaum, “Learning to see physics via visual de-animation,” in NeurIPS , 2017, pp. 153–164
M. Oliu, J. Selva, and S. Escalera, “Folded recurrent neural networks for future video prediction,” in ECCV , 2018, pp. 716–731
2018
Later among the works it cites.
J. Xu, B. Ni, Z. Li, S. Cheng, and X. Yang, “Structure preserving video prediction,” in CVPR , 2018, pp. 1460–1469
2018
Later among the works it cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018, pp. 586–595
2018
Later among the works it cites.
Y. Wang, J. Zhang, H. Zhu, M. Long, J. Wang, and P. S. Yu, “Memory in memory: A predictive neural network for learning higher-order non-stationarity from spatiotemporal dynamics,” in CVPR , 2019, pp. 9154–9162
2019
Later among the works it cites.
Z. Xu, Z. Liu, C. Sun, K. Murphy, W. T. Freeman, J. B. Tenenbaum, and J. Wu, “Unsupervised discovery of parts, structure, and dynamics,” in ICLR , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
C. Finn and S. Levine, “Deep visual foresight for planning robot motion,” in ICRA , 2017, pp. 2786–2793
2017
Cited alongside, same era.
F. Ebert, C. Finn, A. X. Lee, and S. Levine, “Self-supervised visual planning with temporal skip connections,” in CoRL , 2017
2017
Cited alongside, same era.
X. Shi, Z. Gao, L. Lausen, H. Wang, D.-Y. Yeung, W.-k. Wong, and W.-c. Woo, “Deep learning for precipitation nowcasting: A benchmark and a new model,” in NeurIPS , 2017, pp. 5617–5627
2017
Cited alongside, same era.
J. Zhang, Y. Zheng, and D. Qi, “Deep spatio-temporal residual networks for citywide crowd flows prediction.” in AAAI , 2017, pp. 1655–1661
2017
Cited alongside, same era.
P. Bhattacharjee and S. Das, “Temporal coherency based criteria for predicting video frames using deep multi-stage generative adversarial networks,” in NeurIPS , 2017, pp. 4271–4280
2017
Cited alongside, same era.
X. Liang, L. Lee, W. Dai, and E. P. Xing, “Dual motion GAN for future-flow embedded video prediction,” in ICCV , 2017, pp. 1744–1752
2017
Cited alongside, same era.
R. Villegas, J. Yang, Y. Zou, S. Sohn, X. Lin, and H. Lee, “Learning to generate long-term future via hierarchical prediction,” in ICML , 2017, pp. 3560–3569
2017
Cited alongside, same era.
2019
Later among the works it cites.
Y. Wang, L. Jiang, M.-H. Yang, L.-J. Li, M. Long, and L. Fei-Fei, “Eidetic 3D LSTM: A model for video prediction and beyond,” in ICLR , 2019
2019
Later among the works it cites.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in ICML , 2019, pp. 2555–2565
2019
Later among the works it cites.
D. Weissenborn, O. Täckström, and J. Uszkoreit, “Scaling autoregressive video models,” in ICLR , 2019
2019
Later among the works it cites.
T. Kim, S. Ahn, and Y. Bengio, “Variational temporal abstraction,” NeurIPS , vol. 32, pp. 11 570–11 579, 2019
2019
Later among the works it cites.
R. Villegas, A. Pathak, H. Kannan, D. Erhan, Q. V. Le, and H. Lee, “High fidelity video prediction with large stochastic recurrent neural networks,” in NeurIPS , 2019, pp. 81–91
2019
Later among the works it cites.
L. Castrejon, N. Ballas, and A. Courville, “Improved conditional vrnns for video prediction,” in ICCV , 2019, pp. 7608–7617
2019
Later among the works it cites.
K. Greff, R. L. Kaufman, R. Kabra, N. Watters, C. Burgess, D. Zoran, L. Matthey, M. Botvinick, and A. Lerchner, “Multi-object representation learning with iterative variational inference,” in International Conference on Machine Learning . PMLR, 2019, pp. 2424–2433
2019
Later among the works it cites.
IARAI, “Traffic4cast 2019: Traffic map movie forecasting.” https://www.iarai.ac.at/traffic4cast/2019-competition/ , 2019
2019
Later among the works it cites.
W. Yu, Y. Lu, S. Easterbrook, and S. Fidler, “Efficient and information-preserving future frame prediction and beyond,” in ICLR , 2020
2020
Later among the works it cites.
J. Su, W. Byeon, F. Huang, J. Kautz, and A. Anandkumar, “Convolutional tensor-train LSTM for spatio-temporal learning,” in NeurIPS , 2020
2020
Later among the works it cites.
S. Oprea, P. Martinez-Gonzalez, A. Garcia-Garcia, J. A. Castro-Vargas, S. Orts-Escolano, J. Garcia-Rodriguez, and A. Argyros, “A review on deep learning techniques for video prediction,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2020
2020
Later among the works it cites.
M. Kumar, M. Babaeizadeh, D. Erhan, C. Finn, S. Levine, L. Dinh, and D. Kingma, “Videoflow: A flow-based generative model for video,” in ICLR , 2020
2020
Later among the works it cites.
Y. Wu, R. Gao, J. Park, and Q. Chen, “Future video synthesis with object motion prediction,” in CVPR , 2020, pp. 5539–5548
2020
Later among the works it cites.
S. Gur, S. Benaim, and L. Wolf, “Hierarchical patch vae-gan: Generating diverse videos from a single sample,” in NeurIPS , 2020
2020
Later among the works it cites.
J.-Y. Franceschi, E. Delasalles, M. Chen, S. Lamprier, and P. Gallinari, “Stochastic latent residual video prediction,” in ICML , 2020, pp. 3233–3246
2020
Later among the works it cites.
2020
Later among the works it cites.
V. L. Guen and N. Thome, “Disentangling physical dynamics from unknown factors for unsupervised video prediction,” in CVPR , 2020
2020
Later among the works it cites.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” in ICLR , 2020
2020
Later among the works it cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in MICCAI , 2015
2020
Later among the works it cites.
B. Liu, Y. Chen, S. Liu, and H.-S. Kim, “Deep learning in latent space for video prediction and compression,” in CVPR , 2021, pp. 701–710
2021
Closest in time.
B. Wu, S. Nair, R. Martin-Martin, L. Fei-Fei, and C. Finn, “Greedy hierarchical variational autoencoders for large-scale video prediction,” in CVPR , 2021, pp. 2318–2328
2021
Closest in time.
X. Bei, Y. Yang, and S. Soatto, “Learning semantic-aware dynamics for video prediction,” in CVPR , 2021, pp. 902–912
2021
Closest in time.
N. Bodla, G. Shrivastava, R. Chellappa, and A. Shrivastava, “Hierarchical video prediction using relational layouts for human-object interactions,” in CVPR , 2021
2021
Closest in time.
H. Wu, Z. Yao, J. Wang, and M. Long, “Motionrnn: A flexible model for video prediction with spacetime-varying motions,” in CVPR , 2021
2021
Closest in time.