Fetching the paper…
Reading the bibliography…
The ability to predict future visual observations conditioned on past observations and motor commands can enable embodied agents to plan solutions to a variety of tasks in complex environments.
Anticipatory activity of motor cortex neurons in relation to direction of an intended movement
J. Tanji and E. V. Evarts · 1976
Earlier work this paper cites.
An internal model for sensorimotor integration
D. M. Wolpert, Z. Ghahramani, and M. I. Jordan · 1995
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
A tutorial on the cross-entropy method
P. D. Boer, Kroese, S. Mannor, and R. Y. Rubinstein · 2004
Earlier work this paper cites.
A tutorial on the cross-entropy method
P.-T. De Boer, D. P. Kroese, S. Mannor, and R. Y. Rubinstein · 2005
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
The capacity of cognitive control estimated from a perceptual decision making task
T. Wu, A. J. Dufford, M.-A. Mackie, L. J. Egan, and J. Fan · 2016
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
W. Lotter, G. Kreiman, and D. Cox · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson, A. Alahi, and L. Fei-Fei · 2016
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord and O. Vinyals · 2017
Earlier work this paper cites.
Self-supervised visual planning with temporal skip connections
F. Ebert, C. Finn, A. X. Lee, and S. Levine · 2017
Earlier work this paper cites.
To create what you tell: Generating videos from captions
Y. Pan, Z. Qiu, T. Yao, H. Li, and T. Mei · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Earlier work this paper cites.
Video generation from text
Y. Li, M. Min, D. Shen, D. Carlson, and L. Carin · 2018
Earlier work this paper cites.
Imagine this! scripts to compositions to videos
T. Gupta, D. Schwenk, A. Farhadi, D. Hoiem, and A. Kembhavi · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Cited alongside, same era.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2018
Cited alongside, same era.
Stochastic video generation with a learned prior
E. Denton and R. Fergus · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Cited alongside, same era.
Stochastic adversarial video prediction
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Cited alongside, same era.
Towards accurate generative models of video: A new metric & challenges
Scaling autoregressive video models
D. Weissenborn, O. Täckström, and J. Uszkoreit · 2020
Later among the works it cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Later among the works it cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Later among the works it cites.
NUWA: Visual synthesis pre-training for neural visual world creation
C. Wu, J. Liang, L. Ji, F. Yang, Y. Fang, D. Jiang, and N. Duan · 2021
Later among the works it cites.
Greedy hierarchical variational autoencoders for large-scale video prediction
B. Wu, S. Nair, R. Martin-Martin, L. Fei-Fei, and C. Finn · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Cited alongside, same era.
Deep visual mpc-policy learning for navigation
N. Hirose, F. Xia, R. Martín-Martín, A. Sadeghian, and S. Savarese · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Mask-predict: Parallel decoding of conditional masked language models
M. Ghazvininejad, O. Levy, Y. Liu, and L. Zettlemoyer · 2019
Cited alongside, same era.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Cited alongside, same era.
High fidelity video prediction with large stochastic recurrent neural networks
R. Villegas, A. Pathak, H. Kannan, D. Erhan, Q. V. Le, and H. Lee · 2019
Cited alongside, same era.
Adversarial video generation on complex datasets
A. Clark, J. Donahue, and K. Simonyan · 2019
Cited alongside, same era.
M. Babaeizadeh, M. T. Saffar, S. Nair, S. Levine, C. Finn, and D. Erhan · 2021
Later among the works it cites.
Slamp: Stochastic latent appearance and motion prediction
A. K. Akan, E. Erdem, A. Erdem, and F. Güney · 2021
Later among the works it cites.
Stochastic image-to-video synthesis using cinns
M. Dorkenwald, T. Milbich, A. Blattmann, R. Rombach, K. G. Derpanis, and B. Ommer · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2021
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2021
Later among the works it cites.
Maskgit: Masked generative image transformer
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Closest in time.
Stochastic video prediction with structure and motion
A. K. Akan, S. Safadoust, E. Erdem, A. Erdem, and F. Güney · 2022
Closest in time.
Transframer: Arbitrary frame prediction with generative models
C. Nash, J. Carreira, J. Walker, I. Barr, A. Jaegle, M. Malinowski, and P. Battaglia · 2022
Closest in time.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Closest in time.
Masked conditional video diffusion for prediction, generation, and interpolation
V. Voleti, A. Jolicoeur-Martineau, and C. Pal · 2022
Closest in time.
BEit: BERT pre-training of image transformers
H. Bao, L. Dong, S. Piao, and F. Wei · 2022
Closest in time.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Z. Tong, Y. Song, J. Wang, and L. Wang · 2022
Closest in time.
Masked autoencoders as spatiotemporal learners
C. Feichtenhofer, H. Fan, Y. Li, and K. He · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Closest in time.
The unsurprising effectiveness of pre-trained vision models for control
S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta · 2022
Closest in time.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Closest in time.
Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments
S. Srivastava, C. Li, M. Lingelbach, R. Martín-Martín, F. Xia, K. E. Vainio, Z. Lian, C. Gokmen, S. Buch, K. Liu, et al · 2022
Closest in time.