Fetching the paper…
Reading the bibliography…
A video prediction model that generalizes to diverse scenes would enable intelligent agents such as robots to perform a variety of tasks via planning with the model.
S. E. Palmer, “Hierarchical structure in perceptual representation,”
1977
Earlier work this paper cites.
J. J. Verbeek, N. Vlassis, and B. Kröse, “Efficient greedy learning of gaussian mixture models,”
2003
Earlier work this paper cites.
G. E. Hinton, S. Osindero, and Y.-W. Teh, “A fast learning algorithm for deep belief nets,”
2006
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy layer-wise training of deep networks,” in
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
A. massoud Farahmand, A. Shademan, M. Jagersand, and C. Szepesvári, “Model-based and model-free reinforcement learning for visual servoing,” in
2009
Earlier work this paper cites.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, P.-A. Manzagol, and L. Bottou, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion.”
2010
Earlier work this paper cites.
J. Masci, U. Meier, D. Cireşan, and J. Schmidhuber, “Stacked convolutional auto-encoders for hierarchical feature extraction,” in
2011
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,”
2013
Earlier work this paper cites.
B. Boots, A. Byravan, and D. Fox, “Learning predictive models of a depth camera & manipulator from raw execution traces,” in
2014
Earlier work this paper cites.
J. Zhang, S. Shan, M. Kan, and X. Chen, “Coarse-to-fine auto-encoder networks (cfan) for real-time face alignment,” in
2014
Earlier work this paper cites.
V. Kumar, G. C. Nandi, and R. Kala, “Static hand gesture recognition using stacked denoising sparse autoencoders,” in
2014
Earlier work this paper cites.
Y. Qi, Y. Wang, X. Zheng, and Z. Wu, “Robust feature learning by stacked autoencoder with maximum correntropy criterion,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh, “Action-conditional video prediction using deep networks in atari games,” in
2015
Earlier work this paper cites.
J. Walker, A. Gupta, and M. Hebert, “Dense optical flow prediction from a static image,” in
2015
Earlier work this paper cites.
C. K. Sønderby, T. Raiko, L. Maaløe, S. K. Sønderby, and O. Winther, “How to train deep variational autoencoders and probabilistic ladder networks,” in
2016
Earlier work this paper cites.
M. Mathieu, C. Couprie, and Y. LeCun, “Deep multi-scale video prediction beyond mean square error,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Finn, I. Goodfellow, and S. Levine, “Unsupervised learning for physical interaction through video prediction,” in
2016
Earlier work this paper cites.
X. Jia, B. De Brabandere, T. Tuytelaars, and L. V. Gool, “Dynamic filter networks,” in
2016
Earlier work this paper cites.
T. Xue, J. Wu, K. Bouman, and B. Freeman, “Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks,” in
2016
Earlier work this paper cites.
J. Walker, C. Doersch, A. Gupta, and M. Hebert, “An uncertain future: Forecasting from static images using variational autoencoders,” in
2016
Earlier work this paper cites.
R. Shu, J. Brofos, F. Zhang, H. H. Bui, M. Ghavamzadeh, and M. Kochenderfer, “Stochastic video prediction with conditional density estimation,” in
2016
Earlier work this paper cites.
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves
2016
Earlier work this paper cites.
E. P. Ijjina
2016
Earlier work this paper cites.
C. K. Sønderby, T. Raiko, L. Maaløe, S. K. Sønderby, and O. Winther, “Ladder variational autoencoders,” in
2016
Earlier work this paper cites.
S. Zhao, J. Song, and S. Ermon, “Learning hierarchical features from generative models,” in
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in
2016
Earlier work this paper cites.
C. Finn and S. Levine, “Deep visual foresight for planning robot motion,” in
2017
Cited alongside, same era.
N. Kalchbrenner, A. Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu, “Video pixel networks,” in
2017
Cited alongside, same era.
F. Ebert, C. Finn, A. X. Lee, and S. Levine, “Self-supervised visual planning with temporal skip connections,” in
2017
Cited alongside, same era.
A. S. Polydoros and L. Nalpantidis, “Survey of model-based reinforcement learning: Applications on robotics,”
2017
Cited alongside, same era.
X. Liang, L. Lee, W. Dai, and E. P. Xing, “Dual motion gan for future-flow embedded video prediction,” in
2017
Cited alongside, same era.
A. Byravan and D. Fox, “Se3-nets: Learning rigid body motion using deep neural networks,” in
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in
2018
Later among the works it cites.
R. Villegas, A. Pathak, H. Kannan, D. Erhan, Q. V. Le, and H. Lee, “High fidelity video prediction with large stochastic recurrent neural networks,” in
2019
Later among the works it cites.
L. Castrejon, N. Ballas, and A. Courville, “Improved conditional vrnns for video prediction,” in
2019
Later among the works it cites.
A. Xie, F. Ebert, S. Levine, and C. Finn, “Improvisation through physical understanding: Using novel objects as tools with visual foresight,” in
2019
Later among the works it cites.
C. Paxton, Y. Barnoy, K. Katyal, R. Arora, and G. D. Hager, “Visual robot task planning,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
C. Vondrick and A. Torralba, “Generating the future with adversarial transformers,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Z. Liu, R. A. Yeh, X. Tang, Y. Liu, and A. Agarwala, “Video frame synthesis using deep voxel flow,” in
2017
Cited alongside, same era.
B. Chen, W. Wang, and J. Wang, “Video imagination from a single image with transformation generation,” in
2017
Cited alongside, same era.
C. Lu, M. Hirsch, and B. Scholkopf, “Flexible spatio-temporal networks for video prediction,” in
2017
Cited alongside, same era.
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma, “Pixelcnn++: A pixelcnn implementation with discretized logistic mixture likelihood and other modifications,” in
2017
Cited alongside, same era.
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. Johnson, and S. Levine, “Solar: Deep structured representations for model-based reinforcement learning,” in
2019
Later among the works it cites.
N. Hirose, A. Sadeghian, F. Xia, R. Martín-Martín, and S. Savarese, “Vunet: Dynamic scene view synthesis for traversability estimation using an rgb camera,”
2019
Later among the works it cites.
N. Hirose, F. Xia, R. Martín-Martín, A. Sadeghian, and S. Savarese, “Deep visual mpc-policy learning for navigation,”
2019
Later among the works it cites.
Y. Ye, M. Singh, A. Gupta, and S. Tulsiani, “Compositional video prediction,” in
2019
Later among the works it cites.
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn, “Robonet: Large-scale multi-robot learning,” in
2019
Later among the works it cites.
K. Greff, R. L. Kaufman, R. Kabra, N. Watters, C. Burgess, D. Zoran, L. Matthey, M. Botvinick, and A. Lerchner, “Multi-object representation learning with iterative variational inference,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
R. Veerapaneni, J. D. Co-Reyes, M. Chang, M. Janner, C. Finn, J. Wu, J. Tenenbaum, and S. Levine, “Entity abstraction in visual model-based reinforcement learning,” in
2019
Later among the works it cites.
T. Kipf, E. van der Pol, and M. Welling, “Contrastive learning of structured world models,” in
2019
Later among the works it cites.
M. Engelcke, A. R. Kosiorek, O. P. Jones, and I. Posner, “Genesis: Generative scene inference and sampling with object-centric latent representations,” in
2019
Later among the works it cites.
E. Belilovsky, M. Eickenberg, and E. Oyallon, “Greedy layerwise learning can scale to imagenet,” in
2019
Later among the works it cites.
S. Löwe, P. O’Connor, and B. Veeling, “Putting an end to end-to-end: Gradient-isolated learning of representations,” in
2019
Later among the works it cites.
S. Löwe, P. O’Connor, and B. S. Veeling, “Greedy infomax for self-supervised representation learning,” 2019
2019
Later among the works it cites.
L. Maaløe, M. Fraccaro, V. Liévin, and O. Winther, “Biva: A very deep hierarchy of latent variables for generative modeling,” in
2019
Later among the works it cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” 2020
2020
Later among the works it cites.
S. Nair and C. Finn, “Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation,”
2020
Later among the works it cites.
S. Nair, M. Babaeizadeh, C. Finn, S. Levine, and V. Kumar, “Trass: Time reversal as self-supervision,” in
2020
Later among the works it cites.
M. S. Nunes, A. Dehban, P. Moreno, and J. Santos-Victor, “Action-conditioned benchmarking of robotic video prediction models: a comparative study,” in
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Nair, S. Savarese, and C. Finn, “Goal-aware prediction: Learning to model what matters,”
2020
Later among the works it cites.
M. Malinowski, G. Swirszcz, J. Carreira, and V. Patraucean, “Sideways: Depth-parallel training of video models,” in
2020
Later among the works it cites.
A. Vahdat and J. Kautz, “Nvae: A deep hierarchical variational autoencoder,”
2020
Later among the works it cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”
2020
Later among the works it cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,”
2020
Later among the works it cites.