Fetching the paper…
Reading the bibliography…
Learning to predict the long-term future of video frames is notoriously challenging due to inherent ambiguities in the distant future and dramatic amplifications of prediction error through time.
Human activity prediction: Early recognition of ongoing activities from streaming videos
M. S. Ryoo · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Max-margin early event detectors
Minh Hoai and Fernando Torre · 2013
Earlier work this paper cites.
A hierarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
Marc’Aurelio Ranzato, Arthur Szlam, Joan Bruna, Michaël Mathieu, Ronan Collobert, and Sumit Chopra · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, Wei Liu, Yangqing Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Earlier work this paper cites.
Stacked hourglass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Cited alongside, same era.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
Fisher Yu and Vladlen Koltun · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Cited alongside, same era.
Predicting future instance segmentation by forecasting convolutional features
Pauline Luc, Camille Couprie, Yann LeCun, and Jakob Verbeek · 2018
Later among the works it cites.
MoCoGAN: Decomposing motion and content for video generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz · 2018
Later among the works it cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly · 2018
Later among the works it cites.
Video-to-video synthesis
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
Later among the works it cites.
Hierarchical long-term video prediction without supervision
Nevan Wichers, Ruben Villegas, Dumitru Erhan, and Honglak Lee · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dual motion gan for future-flow embedded video prediction
Xiaodan Liang, Lisa Lee, Wei Dai, and Eric P. Xing · 2017
Cited alongside, same era.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2017
Cited alongside, same era.
Predicting deeper into the future of semantic segmentation
Pauline Luc, Natalia Neverova, Camille Couprie, Jacob Verbeek, and Yann LeCun · 2017
Cited alongside, same era.
The pose knows: Video forecasting by generating pose futures
Jacob Walker, Kenneth Marino, Abhinav Gupta, and Martial Hebert · 2017
Cited alongside, same era.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Cited alongside, same era.
Densepose: Dense human pose estimation in the wild
Rıza Alp Güler, Natalia Neverova, and Iasonas Kokkinos · 2018
Cited alongside, same era.
Mt-vae: Learning motion transformations to generate multimodal human dynamics
Xinchen Yan, Akash Rastogi, Ruben Villegas, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Ersin Yumer, and Honglak Lee · 2018
Later among the works it cites.
Pose guided human video generation
Ceyuan Yang, Zhe Wang, Xinge Zhu, Chen Huang, Jianping Shi, and Dahua Lin · 2018
Later among the works it cites.
Bayesian prediction of future street scenes using synthetic likelihoods
Apratim Bhattacharyya, Mario Fritz, and Bernt Schiele · 2019
Later among the works it cites.
Improved conditional vrnns for video prediction
Lluis Castrejon, Nicolas Ballas, and Aaron Courville · 2019
Later among the works it cites.
Efficient video generation on complex datasets
Aidan Clark, Jeff Donahue, and Karen Simonyan · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Later among the works it cites.
Unsupervised keypoint learning for guiding class-conditional video prediction
Yunji Kim, Seonghyeon Nam, In Cho, and Seon Joo Kim · 2019
Later among the works it cites.
Unsupervised learning of object structure and dynamics from videos
Matthias Minderer, Chen Sun, Ruben Villegas, Forrester Cole, Kevin P Murphy, and Honglak Lee · 2019
Later among the works it cites.
Edgeconnect: Structure guided image inpainting using edge prediction
Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Qureshi, and Mehran Ebrahimi · 2019
Later among the works it cites.
High fidelity video prediction with large stochastic recurrent neural networks
Ruben Villegas, Arkanath Pathak, Harini Kannan, Dumitru Erhan, Quoc V Le, and Honglak Lee · 2019
Later among the works it cites.
Improving semantic segmentation via video propagation and label relaxation
Yi Zhu, Karan Sapra, Fitsum A. Reda, Kevin J. Shih, Shawn D. Newsam, Andrew Tao, and Bryan Catanzaro · 2019
Later among the works it cites.