Fetching the paper…
Reading the bibliography…
We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos.
Efficient video generation on complex datasets
A. Clark, J. Donahue, and K. Simonyan · 1907
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 1909
Earlier work this paper cites.
The deepmind jax ecosystem, 2020
I. Babuschkin, K. Baumli, A. Bell, S. Bhupatiraju, J. Bruce, P. Buchlovsky, D. Budden, T. Cai, A. Clark, I. Danihelka, et al · 2010
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. Lewis, and S. Singh · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Recurrent environment simulators
S. Chiappa, S. Racaniere, D. Wierstra, and S. Mohamed · 2017
Earlier work this paper cites.
Video pixel networks
N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
W. Lotter, G. Kreiman, and D. Cox · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Neural scene representation and rendering
S. M. A. Eslami, D. J. Rezende, F. Besse, F. Viola, A. S. Morcos, M. Garnelo, A. Ruderman, A. A. Rusu, I. Danihelka, K. Gregor, D. P. Reichert, L. Buesing, T. Weber, O. Vinyals, D. Rosenbaum, N. Rabinowitz, H. King, C. Hillier, M. Botvinick, D. Wierstra, K. Kavukcuoglu, and D. Hassabis · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Recurrent experience replay in distributed reinforcement learning
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Procedural content generation via machine learning (PCGML)
A. Summerville, S. Snodgrass, M. Guzdial, C. Holmgård, A. K. Hoover, A. Isaksen, A. Nealen, and J. Togelius · 2018
Earlier work this paper cites.
Behavioral cloning from observation
F. Torabi, G. Warnell, and P. Stone · 2018
Earlier work this paper cites.
J. Clune · 2019
Earlier work this paper cites.
Imitating latent policies from observation
A. Edwards, H. Sahni, Y. Schroecker, and C. Isbell · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Learning what you can do before doing anything
O. Rybkin*, K. Pertsch*, K. G. Derpanis, K. Daniilidis, and A. Jaegle · 2019
Earlier work this paper cites.
FVD: A new metric for video generation, 2019
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2019
Earlier work this paper cites.
Neural game engine: Accurate learning ofgeneralizable forward models from pixels
C. Bamford and S. M. Lucas · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman · 2020
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Earlier work this paper cites.
Query-key normalization for transformers
A. Henry, P. R. Dachapally, S. S. Pawar, and Y. Chen · 2020
Cited alongside, same era.
A domain-specific supercomputer for training deep neural networks
N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson · 2020
Cited alongside, same era.
Learning to simulate dynamic environments with gamegan
S. W. Kim, Y. Zhou, J. Philion, A. Torralba, and S. Fidler · 2020
Cited alongside, same era.
Transformation-based adversarial video prediction on large-scale data
P. Luc, A. Clark, S. Dieleman, D. de Las Casas, Y. Doron, A. Cassirer, and K. Simonyan · 2020
Cited alongside, same era.
Action-conditioned benchmarking of robotic video prediction models: a comparative study
M. S. Nunes, A. Dehban, P. Moreno, and J. Santos-Victor · 2020
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, R. Gontijo-Lopes, B. K. Ayan, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi · 2022
Later among the works it cites.
Nüwa: Visual synthesis pre-training for neural visual world creation
C. Wu, J. Liang, L. Ji, F. Yang, Y. Fang, D. Jiang, and N. Duan · 2022
Later among the works it cites.
Become a proficient player with limited data through watching pure videos
W. Ye, Y. Zhang, P. Abbeel, and Y. Gao · 2022
Later among the works it cites.
Human-timescale adaptation in an open-ended task space
J. Bauer, K. Baumli, F. Behbahani, A. Bhoopchand, N. Bradley-Schmieg, M. Chang, N. Clay, A. Collister, V. Dasagi, L. Gonzalez, K. Gregor, E. Hughes, S. Kashem, M. Loks-Thompson, H. Openshaw, J. Parker-Holder, S. Pathak, N. Perez-Nieves, N. Rakicevic, T. Rocktäschel, Y. Schroecker, S. Singh, J. Sygnowski, K. Tuyls, S. York, A. Zacherl, and L. M. Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Spatial-temporal transformer networks for traffic flow forecasting
M. Xu, W. Dai, C. Liu, X. Gao, W. Lin, G.-J. Qi, and H. Xiong · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Mastering atari with discrete world models
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba · 2021
Cited alongside, same era.
Drivegan: Towards a controllable high-quality neural simulation
S. W. Kim, J. Philion, A. Torralba, and S. Fidler · 2021
Cited alongside, same era.
Ccvs: Context-aware controllable video synthesis
G. Le Moing, J. Ponce, and C. Schmid · 2021
Cited alongside, same era.
Align your latents: High-resolution video synthesis with latent diffusion models
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis · 2023
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. Riquelme Ruiz, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. V. Steenkiste, G. F. Elsayed, A. Mahendran, F. Yu, A. Oliver, F. Huot, J. Bastings, M. Collier, A. A. Gritsenko, V. Birodkar, C. N. Vasconcelos, Y. Tay, T. Mensink, A. Kolesnikov, F. Pavetic, D. Tran, T. Kipf, M. Lucic, X. Zhai, D. Keysers, J. J. Harmsen, and N. Houlsby · 2023
Later among the works it cites.
Structure and content-guided video synthesis with diffusion models
P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis · 2023
Later among the works it cites.
Maskvit: Masked visual pre-training for video prediction
A. Gupta, S. Tian, Y. Zhang, J. Wu, R. Martín-Martín, and L. Fei-Fei · 2023
Later among the works it cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2023
Later among the works it cites.
Gaia-1: A generative world model for autonomous driving, 2023
A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado · 2023
Later among the works it cites.
Transformers are sample-efficient world models
V. Micheli, E. Alonso, and F. Fleuret · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Later among the works it cites.
Transformer-based world models are happy with 100k interactions
J. Robine, M. Höftmann, T. Uelwer, and S. Harmeling · 2023
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman · 2023
Later among the works it cites.
Prompt-guided level generation
S. Sudhakaran, M. González-Duque, C. Glanois, M. Freiberger, E. Najarro, and S. Risi · 2023
Later among the works it cites.
Level generation through large language models
G. Todd, S. Earle, M. U. Nasir, M. C. Green, and J. Togelius · 2023
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual descriptions
R. Villegas, M. Babaeizadeh, P.-J. Kindermans, H. Moraldo, H. Zhang, M. T. Saffar, S. Castro, J. Kunze, and D. Erhan · 2023
Later among the works it cites.
Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2023
Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Chen, Y. Wang, P. Luo, Z. Liu, Y. Wang, L. Wang, and Y. Qiao · 2023
Later among the works it cites.
From word models to world models: Translating from natural language to the probabilistic language of thought, 2023
L. Wong, G. Grand, A. K. Lew, N. D. Goodman, V. K. Mansinghka, J. Andreas, and J. B. Tenenbaum · 2023
Later among the works it cites.
Temporally consistent transformers for video generation
W. Yan, D. Hafner, S. James, and P. Abbeel · 2023
Later among the works it cites.
Learning interactive real-world simulators
M. Yang, Y. Du, K. Ghasemipour, J. Tompson, D. Schuurmans, and P. Abbeel · 2023
Later among the works it cites.
Magvit: Masked generative video transformer
L. Yu, Y. Cheng, K. Sohn, J. Lezama, H. Zhang, H. Chang, A. G. Hauptmann, M. Yang, Y. Hao, I. Essa, and L. Jiang · 2023
Later among the works it cites.
Video generation models as world simulators
T. Brooks, B. Peebles, C. Homes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. W. Y. Ng, R. Wang, and A. Ramesh · 2024
Closest in time.
Learning to act without actions
D. Schmidt and M. Jiang · 2024
Closest in time.