Fetching the paper…
Reading the bibliography…
The task of video prediction and generation is known to be notoriously difficult, with the research in this area largely limited to short-term predictions.
Recognizing human actions: A local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Generative adversarial networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L. Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Semi-supervised learning with ladder networks
Antti Rasmus, Mathias Berglund, M. Honkala, H. Valpola, and T. Raiko · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Deep unsupervised clustering with gaussian mixture variational autoencoders
Nat Dilokthanakul, Pedro A. M. Mediano, Marta Garnelo, M. J. Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2016
Earlier work this paper cites.
Ladder variational autoencoders
Casper Kaae Sønderby, Tapani Raiko, Lars Maaløe, Søren Kaae Sønderby, and Ole Winther · 2016
Earlier work this paper cites.
Se3-nets: Learning rigid body motion using deep neural networks
Arunkumar Byravan and Dieter Fox · 2017
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2017
Earlier work this paper cites.
Variational deep embedding: An unsupervised and generative approach to clustering
Zhuxi Jiang, Yin Zheng, Huachun Tan, Bangsheng Tang, and Hanning Zhou · 2017
Earlier work this paper cites.
Video pixel networks
Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Fast-slow recurrent neural networks
Asier Mujika, Florian Meier, and Angelika Steger · 2017
Earlier work this paper cites.
Parallel multiscale autoregressive density estimation
Scott Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gómez Colmenarejo, Ziyu Wang, Dan Belov, and Nando de Freitas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Generating the future with adversarial transformers
Carl Vondrick and Antonio Torralba · 2017
Cited alongside, same era.
Stochastic video generation with a learned prior
Emily L. Denton and Rob Fergus · 2018
Cited alongside, same era.
Neural scene representation and rendering
S. M. Ali Eslami, Danilo Jimenez Rezende, Frederic Besse, Fabio Viola, Ari S. Morcos, Marta Garnelo, Avraham Ruderman, Andrei A. Rusu, Ivo Danihelka, Karol Gregor, David P. Reichert, Lars Buesing, Theophane Weber, Oriol Vinyals, Dan Rosenbaum, Neil Rabinowitz, Helen King, Chloe Hillier, Matt Botvinick, Daan Wierstra, Koray Kavukcuoglu, and Demis Hassabis · 2018
Cited alongside, same era.
Stochastic adversarial video prediction
Alex X. Lee, Richard Zhang, Frederik Ebert, Pieter Abbeel, Chelsea Finn, and Sergey Levine · 2018
Cited alongside, same era.
Latent video transformer
Ruslan Rakhimov, Denis Volkhonskiy, Alexey Artemov, Denis Zorin, and Evgeny Burnaev · 2020
Later among the works it cites.
Nvae: A deep hierarchical variational autoencoder
Arash Vahdat and J. Kautz · 2020
Later among the works it cites.
Fitvid: Overfitting in pixel-level video prediction
Mohammad Babaeizadeh, Mohammad Taghi Saffar, Suraj Nair, Sergey Levine, Chelsea Finn, and Dumitru Erhan · 2021
Later among the works it cites.
Very deep vaes generalize autoregressive models and can outperform them on images
Rewon Child · 2021
Later among the works it cites.
Multi-facet clustering variational autoencoders
Fabian Falck, Haoting Zhang, Matthew Willetts, George Nicholson, Christopher Yau, and Christopher C. Holmes · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive skip intervals: Temporal abstraction for recurrent dynamical models
Alexander Neitz, Giambattista Parascandolo, Stefan Bauer, and Bernhard Schölkopf · 2018
Cited alongside, same era.
Improved conditional vrnns for video prediction
Lluís Castrejón, Nicolas Ballas, and Aaron C. Courville · 2019
Cited alongside, same era.
Adversarial video generation on complex datasets
Aidan Clark, Jeff Donahue, and Karen Simonyan · 2019
Cited alongside, same era.
Hierarchical generative modeling for controllable speech synthesis
Wei-Ning Hsu, Y. Zhang, Ron J. Weiss, H. Zen, Yonghui Wu, Yuxuan Wang, Yuan Cao, Ye Jia, Z. Chen, Jonathan Shen, P. Nguyen, and Ruoming Pang · 2019
Cited alongside, same era.
Time-agnostic prediction: Predicting predictable video frames
Dinesh Jayaraman, Frederik Ebert, Alexei Efros, and Sergey Levine · 2019
Cited alongside, same era.
Variational temporal abstraction
Taesup Kim, Sungjin Ahn, and Yoshua Bengio · 2019
Cited alongside, same era.
Compile: Compositional imitation learning and execution
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2019
Cited alongside, same era.
Arrowgan : Learning to generate videos by learning arrow of time
Kibeom Hong, Youngjung Uh, and Hyeran Byun · 2021
Later among the works it cites.
Video prediction recalling long-term motion context via memory alignment learning
Sangmin Lee, Hak Gu Kim, Dae Hwi Choi, Hyung-Il Kim, and Yong Man Ro · 2021
Later among the works it cites.
Clockwork variational autoencoders
Vaibhav Saxena, Jimmy Ba, and Danijar Hafner · 2021
Later among the works it cites.
Greedy hierarchical variational autoencoders for large-scale video prediction
Bohan Wu, Suraj Nair, Roberto Martín-Martín, Li Fei-Fei, and Chelsea Finn · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas · 2021
Later among the works it cites.
Episodic memory for subjective-timescale models
Alexey Zakharov, Matthew Crosby, and Zafeirios Fountas · 2021
Later among the works it cites.
A Predictive Processing Model of Episodic Memory and Time Perception
Zafeirios Fountas, Anastasia Sylaidi, Kyriacos Nikiforou, Anil K. Seth, Murray Shanahan, and Warrick Roseboom · 2022
Closest in time.
Flexible diffusion modeling of long videos
William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Dietrich Weilbach, and Frank Wood · 2022
Closest in time.
Beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2022
Closest in time.
Diffusion models for video prediction and infilling
Tobias Höppe, Arash Mehrjou, Stefan Bauer, Didrik Nielsen, and Andrea Dittadi · 2022
Closest in time.
Diffusion probabilistic modeling for video generation
Ruihan Yang, Prakhar Srivastava, and Stephan Mandt · 2022
Closest in time.
Variational predictive routing with nested subjective timescales
Alexey Zakharov, Qinghai Guo, and Zafeirios Fountas · 2022
Closest in time.