Fetching the paper…
Reading the bibliography…
Much of recent research has been devoted to video prediction and generation, yet most of the previous works have demonstrated only limited success in generating videos on short-term horizons.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The recurrent temporal restricted Boltzmann machine
I. Sutskever, G. E. Hinton, and G. W. Taylor · 2009
Earlier work this paper cites.
Latent structured models for human pose estimation
C. Ionescu, F. Li, and C. Sminchisescu · 2011
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Modeling deep temporal dependencies with recurrent “grammar cells”
V. Michalski, R. Memisevic, and K. Konda · 2014
Earlier work this paper cites.
Structured recurrent temporal restricted Boltzmann machines
R. Mittelman, B. Kuipers, S. Savarese, and H. Lee · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Cited alongside, same era.
Learning to linearize under uncertainty
R. Goroshin, M. Mathieu, and Y. LeCun · 2015
Cited alongside, same era.
Learning image representations tied to ego-motion
D. Jayaraman and K. Grauman · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in Atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Cited alongside, same era.
Deep visual analogy-making
S. E. Reed, Y. Zhang, Y. Zhang, and H. Lee · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion
D. Jayaraman and K. Grauman · 2016
Later among the works it cites.
Deep multi-scale video prediction beyond mean square error
M. Mathieu, C. Couprie, and Y. LeCun · 2016
Later among the works it cites.
Improved training of wasserstein GANs
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville · 2017
Later among the works it cites.
MobileNets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Later among the works it cites.
Video pixel networks
N. Kalchbrenner, A. v. d. Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu · 2017
Later among the works it cites.
Deep predictive coding networks for video prediction and unsupervised learning
W. Lotter, G. Kreiman, and D. Cox · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Cited alongside, same era.
Decomposing motion and content for natural video sequence prediction
R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee
Cited in the paper.
Learning to generate long-term future via hierarchical prediction
R. Villegas, J. Yang, Y. Zou, S. Sohn, X. Lin, and H. Lee
Cited in the paper.
Later among the works it cites.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2018
Closest in time.
Stochastic video generation with a learned prior
E. Denton and R. Fergus · 2018
Closest in time.