Fetching the paper…
Reading the bibliography…
Extracting and predicting object structure and dynamics from videos without supervision is a major challenge in machine learning.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
A Recurrent Latent Variable Model for Sequential Data
J. Chung, K. Kastner, L. Dinh, K. Goel, A. Courville, and Y. Bengio · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Earlier work this paper cites.
Unsupervised Learning of Video Representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Earlier work this paper cites.
Generating Sentences from a Continuous Space
S. R. Bowman, L. Vilnis, O. Vinyals, A. M. Dai, R. Jozefowicz, and S. Bengio · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
M. Mathieu, C. Couprie, and Y. LeCun · 2016
Earlier work this paper cites.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
T. Xue, J. Wu, K. Bouman, and B. Freeman · 2016
Earlier work this paper cites.
Unsupervised Learning of Disentangled Representations from Video
E. Denton and V. Birodkar · 2017
Cited alongside, same era.
Google vizier: A service for black-box optimization
D. Golovin, B. Solnik, S. Moitra, G. Kochanski, J. Karro, and D. Sculley · 2017
Cited alongside, same era.
Deep predictive coding networks for video prediction and unsupervised learning
W. Lotter, G. Kreiman, and D. Cox · 2017
Cited alongside, same era.
Fixing a Broken ELBO
A. A. Alemi, B. Poole, I. Fischer, J. V. Dillon, R. A. Saurous, and K. Murphy · 2018
Cited alongside, same era.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, R. Erhan, Dumitru an Campbell, and S. Levine · 2018
Cited alongside, same era.
Accurate and Diverse Sampling of Sequences based on a "Best of Many" Sample Objective
A. Bhattacharyya, B. Schiele, and M. Fritz · 2018
Cited alongside, same era.
Towards Accurate Generative Models of Video: A New Metric & Challenges
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Later among the works it cites.
The Pose Knows: Video Forecasting by Generating Pose Futures
J. Walker, K. Marino, A. Gupta, and M. Hebert · 2018
Later among the works it cites.
Video-to-Video Synthesis
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, G. Liu, A. Tao, J. Kautz, and B. Catanzaro · 2018
Later among the works it cites.
Hierarchical Long-term Video Prediction without Supervision
N. Wichers, R. Villegas, D. Erhan, and H. Lee · 2018
Later among the works it cites.
Mt-vae: Learning motion transformations to generate multimodal human dynamics
X. Yan, A. Rastogi, R. Villegas, K. Sunkavalli, E. Shechtman, S. Hadap, E. Yumer, and H. Lee · 2018
Later among the works it cites.
Unsupervised Discovery of Object Landmarks as Structural Representations
Y. Zhang, Y. Guo, Y. Jin, Y. Luo, Z. He, and H. Lee · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Chan, S. Ginosar, T. Zhou, and A. A. Efros · 2018
Cited alongside, same era.
Stochastic Video Generation with a Learned Prior
E. Denton and R. Fergus · 2018
Cited alongside, same era.
Conditional Image Generation for Learning the Structure of Visual Objects
T. Jakab, A. Gupta, H. Bilen, and A. Vedaldi · 2018
Cited alongside, same era.
Stochastic Adversarial Video Prediction
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Cited alongside, same era.
DeepMind Control Suite
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. Lillicrap, and M. Riedmiller · 2018
Cited alongside, same era.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Cited alongside, same era.
Later among the works it cites.
Learning Character-Agnostic Motion for Motion Retargeting in 2D
K. Aberman, R. Wu, D. Lischinski, B. Chen, and D. Cohen-Or · 2019
Closest in time.
Learning Latent Dynamics for Planning from Pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Closest in time.
Unsupervised learning of object keypoints for perception and control
T. Kulkarni, A. Gupta, C. Ionescu, S. Borgeaud, M. Reynolds, A. Zisserman, and V. Mnih · 2019
Closest in time.
Stochastic Prediction of Multi-Agent Interactions From Partial Observations
C. Sun, P. Karlsson, J. Wu, J. B. Tenenbaum, and K. Murphy · 2019
Closest in time.
Unsupervised Discovery of Parts, Structure, and Dynamics
Z. Xu, Z. Liu, C. Sun, K. Murphy, W. T. Freeman, J. B. Tenenbaum, and J. Wu · 2019
Closest in time.
Generating Multi-Agent Trajectories using Programmatic Weak Supervision
E. Zhan, S. Zheng, Y. Yue, L. Sha, and P. Lucey · 2019
Closest in time.