Fetching the paper…
Reading the bibliography…
Autonomous systems not only need to understand their current environment, but should also be able to predict future actions conditioned on past states, for instance based on captured camera frames.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah Hihi and Yoshua Bengio · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun, Corinna Cortes, and Christopher JC Burges · 1998
Earlier work this paper cites.
Recognizing human actions: A local SVM approach
Christian Schuldt, Ivan Laptev, and Barbara Caputo · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2004
Earlier work this paper cites.
Modec: Multimodal decomposable models for human pose estimation
Ben Sapp and Ben Taskar · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
A clockwork RNN
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber · 2014
Earlier work this paper cites.
Modeling deep temporal dependencies with recurrent grammar cells
Vincent Michalski, Roland Memisevic, and Kishore Konda · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using LSTMs
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Cited alongside, same era.
Dynamic filter networks
Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool · 2016
Cited alongside, same era.
Spatio-temporal video autoencoder with differentiable memory
Viorica Pătrăucean, Ankur Handa, and Roberto Cipolla · 2016
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2016
Cited alongside, same era.
Ladder variational autoencoders
Casper Kaae Sønderby, Tapani Raiko, Lars Maaløe, Søren Kaae Sønderby, and Ole Winther · 2016
Cited alongside, same era.
Realtime multi-person 2D pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Cited alongside, same era.
Predicting future instance segmentation by forecasting convolutional features
Pauline Luc, Camille Couprie, Yann Lecun, and Jakob Verbeek · 2018
Later among the works it cites.
Improved conditional VRNNs for video prediction
Lluis Castrejon, Nicolas Ballas, and Aaron Courville · 2019
Later among the works it cites.
Video generation from single semantic label map
Junting Pan, Chengyu Wang, Xu Jia, Jing Shao, Lu Sheng, Junjie Yan, and Xiaogang Wang · 2019
Later among the works it cites.
Image quality assessment through FSIM, SSIM, MSE and PSNR—a comparative study
Umme Sara, Morium Akter, and Mohammad Shorif Uddin · 2019
Later among the works it cites.
High fidelity video prediction with large stochastic recurrent neural networks
Ruben Villegas, Arkanath Pathak, Harini Kannan, Dumitru Erhan, Quoc V Le, and Honglak Lee · 2019
Later among the works it cites.
Motion segmentation using frequency domain transformer networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Recurrent ladder networks
Isabeau Prémont-Schwarz, Alexander Ilin, Tele Hao, Antti Rasmus, Rinu Boney, and Harri Valpola · 2017
Cited alongside, same era.
Deep learning for precipitation nowcasting: A benchmark and a new model
Xingjian Shi, Zhihan Gao, Leonard Lausen, Hao Wang, Dit-Yan Yeung, Wai-kin Wong, and Wang-chun Woo · 2017
Cited alongside, same era.
Learning to generate long-term future via hierarchical prediction
Ruben Villegas, Jimei Yang, Yuliang Zou, Sungryull Sohn, Xunyu Lin, and Honglak Lee · 2017
Cited alongside, same era.
PredRNN: Recurrent neural networks for predictive learning using spatiotemporal LSTMs
Yunbo Wang, Mingsheng Long, Jianmin Wang, Zhifeng Gao, and Philip S Yu · 2017
Cited alongside, same era.
Hafez Farazi and Sven Behnke · 2020
Later among the works it cites.
Long-term human video generation of multiple futures using poses
Naoya Fushishita, Antonio Tejero-de Pablos, Yusuke Mukuta, and Tatsuya Harada · 2020
Later among the works it cites.
Disentangling physical dynamics from unknown factors for unsupervised video prediction
Vincent Le Guen and Nicolas Thome · 2020
Later among the works it cites.
A review on deep learning techniques for video prediction
Sergiu Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts-Escolano, Jose Garcia-Rodriguez, and Antonis Argyros · 2020
Later among the works it cites.
SLAMP: Stochastic latent appearance and motion prediction
Adil Kaan Akan, Erkut Erdem, Aykut Erdem, and Fatma Güney · 2021
Later among the works it cites.
Local frequency domain transformer networks for video prediction
Hafez Farazi, Jan Nogga, and Sven Behnke · 2021
Later among the works it cites.
SynPick: A dataset for dynamic bin picking scene understanding
Arul Selvam Periyasamy, Max Schwarz, and Sven Behnke · 2021
Later among the works it cites.
Clockwork variational autoencoders
Vaibhav Saxena, Jimmy Ba, and Danijar Hafner · 2021
Later among the works it cites.
Greedy hierarchical variational autoencoders for large-scale video prediction
Bohan Wu, Suraj Nair, Roberto Martin-Martin, Li Fei-Fei, and Chelsea Finn · 2021
Later among the works it cites.
PredRNN: A recurrent neural network for spatiotemporal predictive learning
Yunbo Wang, Haixu Wu, Jianjin Zhang, Zhifeng Gao, Jianmin Wang, Philip Yu, and Mingsheng Long · 2022
Closest in time.