Fetching the paper…
Reading the bibliography…
We propose a hierarchical approach for making long-term predictions of future frames.
Two-frame motion estimation based on polynomial expansion
Farneback, G · 2003
Earlier work this paper cites.
The recurrent temporal restricted boltzmann machine
Sutskever, I., Hinton, G. E., and Taylor, G. W · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
From actemes to action: A strongly-supervised representation for detailed action understanding
Weiyu Zhang, M. Z. and Derpanis, K · 2013
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Ionescu, C., Papava, D., Olaru, V., and Sminchisescu, C · 2014
Earlier work this paper cites.
Modeling deep temporal dependencies with recurrent "grammar cells"
Michalski, V., Memisevic, R., and Konda, K · 2014
Earlier work this paper cites.
Structured recurrent temporal restricted boltzmann machines
Mittelman, R., Kuipers, B., Savarese, S., and Lee, H · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
Ranzato, M., Szlam, A., Bruna, J., Mathieu, M., Collobert, R., and Chopra, S · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Patch to the future: Unsupervised visual prediction
Walker, J., Gupta , A., and Hebert , M · 2014
Earlier work this paper cites.
Recurrent network models for human dynamics
Fragkiadaki, K., Levine, S., Felsen, P., and Malik, J · 2015
Earlier work this paper cites.
Learning to linearize under uncertainty
Goroshin, R., Mathieu, M., and LeCun, Y · 2015
Cited alongside, same era.
Learning image representations tied to ego-motion
Jayaraman, D. and Grauman, K · 2015
Cited alongside, same era.
Modeling of Dynamic Environments for Visual Forecasting of American Football Plays
Lee, N · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Cited alongside, same era.
Deep visual analogy-making
Reed, S., Zhang, Y., Zhang, Y., and Lee, H · 2015
Cited alongside, same era.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Shi, X., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-k., and WOO, W.-c · 2015
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
Finn, C., Goodfellow, I. J., and Levine, S · 2016
Later among the works it cites.
Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion
Jayaraman, D. and Grauman, K · 2016
Later among the works it cites.
Deep multi-scale video prediction beyond mean square error
Mathieu, M., Couprie, C., and LeCun, Y · 2016
Later among the works it cites.
Stacked hourglass networks for human pose estimation
Newell, A., Yang, K., and Deng, J · 2016
Later among the works it cites.
Generating videos with scene dynamics
Vondrick, C., Pirsiavash, H., and Torralba, A · 2016
Later among the works it cites.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
Xue, T., Wu, J., Bouman, K. L., and Freeman, W. T · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
Srivastava, N., Mansimov, E., and Salakhudinov, R · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Generating images with perceptual similarity metrics based on deep networks
Dosovitskiy, A. and Brox, T · 2016
Cited alongside, same era.
Learning what and where to draw
Reed, S., Akata, Z., Mohan, S., Tenka, S., Schiele, B., and Lee, H
Cited in the paper.
Generative adversarial text-to-image synthesis
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H
Cited in the paper.
Later among the works it cites.
Forecasting human dynamics from static images
Chao, Y.-W., Yang, J., Price, B., Cohen, S., and Deng, J · 2017
Closest in time.
Deep predictive coding networks for video prediction and unsupervised learning
Lotter, W., Kreiman, G., and Cox, D · 2017
Closest in time.
Decomposing motion and content for natural video sequence prediction
Villegas, R., Yang, J., Hong, S., Lin, X., and Lee, H · 2017
Closest in time.
A data-driven approach for event prediction
Yuen, J. and Torralba, A · 2017
Closest in time.