Fetching the paper…
Reading the bibliography…
We propose a probabilistic video model, the Video Pixel Network (VPN), that estimates the discrete joint distribution of the raw pixel values in a video.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom · 2013
Earlier work this paper cites.
Semantic image segmentation with deep convolutional nets and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
Marc’Aurelio Ranzato, Arthur Szlam, Joan Bruna, Michaël Mathieu, Ronan Collobert, and Sumit Chopra · 2014
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michaël Mathieu, Camille Couprie, and Yann LeCun · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L. Lewis, and Satinder P. Singh · 2015
Cited alongside, same era.
Spatio-temporal video autoencoder with differentiable memory
Viorica Patraucean, Ankur Handa, and Roberto Cipolla · 2015
Cited alongside, same era.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
Fisher Yu and Vladlen Koltun · 2015
Cited alongside, same era.
Bert De Brabandere, Xu Jia, Tinne Tuytelaars, and Luc Van Gool · 2016
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian J. Goodfellow, and Sergey Levine · 2016
Closest in time.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Closest in time.
Grid long short-term memory
Nal Kalchbrenner, Ivo Danihelka, and Alex Graves · 2016
Closest in time.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov
Cited in the paper.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber
Cited in the paper.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu
Cited in the paper.
Pixel recurrent neural networks
Aäron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu
Cited in the paper.
Conditional image generation with pixelcnn decoders
Aäron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu
Cited in the paper.