Fetching the paper…
Reading the bibliography…
World models aim to learn action-controlled future prediction and have proven essential for the development of intelligent agents.
Dyna, an Integrated Architecture for Learning, Planning, and Reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Premotor Cortex and the Recognition of Motor Actions
Rizzolatti, G., Fadiga, L., Gallese, V., and Fogassi, L · 1996
Earlier work this paper cites.
Generalization in Vision and Motor Control
Poggio, T. and Bizzi, E · 2004
Earlier work this paper cites.
Neuronal Correlates of a Perceptual Decision in Ventral Premotor Cortex
Romo, R., Hernández, A., and Zainos, A · 2004
Earlier work this paper cites.
A Tutorial on the Cross-Entropy Method
De Boer, P.-T., Kroese, D. P., Mannor, S., and Rubinstein, R. Y · 2005
Earlier work this paper cites.
Image Quality Metrics: PSNR vs. SSIM
Hore, A. and Ziou, D · 2010
Earlier work this paper cites.
Locomotor Primitives in Newborn Babies and Their Development
Dominici, N., Ivanenko, Y. P., Cappellini, G., d’Avella, A., Mondì, V., Cicchese, M., Fabiano, A., Silei, T., Di Paolo, A., Giannini, C., et al · 2011
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P. K. and Welling, M · 2014
Earlier work this paper cites.
Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M · 2015
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al · 2016
Earlier work this paper cites.
Aggressive Driving with Model Predictive Path Integral Control
Williams, G., Drews, P., Goldfain, B., Rehg, J. M., and Theodorou, E. A · 2016
Earlier work this paper cites.
Deep Variational Information Bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2017
Earlier work this paper cites.
Understanding Disentangling in β \beta -VAE
Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A · 2017
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Self-Supervised Visual Planning with Temporal Skip Connections
Ebert, F., Finn, C., Lee, A. X., and Levine, S · 2017
Earlier work this paper cites.
The “Something Something” Video Database for Learning and Evaluating Visual Common Sense
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al · 2017
Earlier work this paper cites.
beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework
Higgins, I., Matthey, L., Pal, A., Burgess, C. P., Glorot, X., Botvinick, M. M., Mohamed, S., and Lerchner, A · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Attention is All You Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Earlier work this paper cites.
Recurrent World Models Facilitate Policy Evolution
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
McInnes, L., Healy, J., and Melville, J · 2018
Earlier work this paper cites.
Gotta Learn Fast: A New Benchmark for Generalization in RL
Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Towards Accurate Generative Models of Video: A New Metric & Challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Imitating Latent Policies from Observation
Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Learning What You Can Do before Doing Anything
Rybkin, O., Pertsch, K., Derpanis, K. G., Daniilidis, K., and Jaegle, A · 2019
Earlier work this paper cites.
Habitat: A Platform for Embodied AI Research
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al · 2019
Earlier work this paper cites.
nuScenes: A Multimodal Dataset for Autonomous Driving
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O · 2020
Earlier work this paper cites.
Leveraging Procedural Generation to Benchmark Reinforcement Learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Earlier work this paper cites.
Model-Based Reinforcement Learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2020
Earlier work this paper cites.
Learning to Simulate Dynamic Environments with GameGAN
Kim, S. W., Zhou, Y., Philion, J., Torralba, A., and Fidler, S · 2020
Earlier work this paper cites.
Deep Dynamics Models for Learning Dexterous Manipulation
Nagabandi, A., Konolige, K., Levine, S., and Kumar, V · 2020
Earlier work this paper cites.
Learning Predictive Models From Observation and Interaction
Schmeckpeper, K., Xie, A., Rybkin, O., Tian, S., Daniilidis, K., Levine, S., and Finn, C · 2020
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
RoboDesk: A Multi-Task Reinforcement Learning Benchmark
Kannan, H., Hafner, D., Finn, C., and Erhan, D · 2021
Earlier work this paper cites.
DriveGAN: Towards a Controllable High-Quality Neural Simulation
Kim, S. W., Philion, J., Torralba, A., and Fidler, S · 2021
Earlier work this paper cites.
Playable Video Generation
Menapace, W., Lathuiliere, S., Tulyakov, S., Siarohin, A., and Ricci, E · 2021
Cited alongside, same era.
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos
Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Cited alongside, same era.
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Cited alongside, same era.
Latent Video Diffusion Models for High-Fidelity Long Video Generation
He, Y., Yang, T., Zhang, Y., Shan, Y., and Chen, Q · 2022
Cited alongside, same era.
Elucidating the Design Space of Diffusion-Based Generative Models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
GenEx: Generating an Explorable World
Lu, T., Shu, T., Xiao, J., Ye, L., Wang, J., Peng, C., Wei, C., Khashabi, D., Chellappa, R., Yuille, A., et al · 2024
Later among the works it cites.
GenRL: Multimodal-Foundation World Models for Generalization in Embodied Agents
Mazzaglia, P., Verbelen, T., Dhoedt, B., Courville, A., and Rajeswar, S · 2024
Later among the works it cites.
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
O’Neill, A., Rehman, A., Gupta, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., et al · 2024
Later among the works it cites.
Scaling Laws for Pre-Training Agents and World Models
Pearce, T., Rashid, T., Bignell, D., Georgescu, R., Devlin, S., and Hofmann, K · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., et al · 2022
Cited alongside, same era.
Playable Environments: Video Manipulation in Space and Time
Menapace, W., Lathuilière, S., Siarohin, A., Theobalt, C., Tulyakov, S., Golyanik, V., and Ricci, E · 2022
Cited alongside, same era.
A Generalist Agent
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Cited alongside, same era.
Reinforcement Learning with Action-Free Pre-Training from Videos
Seo, Y., Lee, K., James, S. L., and Abbeel, P · 2022
Cited alongside, same era.
DayDreamer: World Models for Physical Robot Learning
Wu, P., Escontrela, A., Hafner, D., Abbeel, P., and Goldberg, K · 2022
Cited alongside, same era.
SelfD: Self-Learning Large-Scale Driving Policies From the Web
Zhang, J., Zhu, R., and Ohn-Bar, E · 2022
Cited alongside, same era.
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
Zhu, Y., Wong, J., Mandlekar, A., Martín-Martín, R., Joshi, A., Nasiriany, S., and Zhu, Y · 2022
Cited alongside, same era.
Raad, M. A., Ahuja, A., Barros, C., Besse, F., Bolt, A., Bolton, A., Brownfield, B., Buttimore, G., Cant, M., Chakera, S., et al · 2024
Later among the works it cites.
AVID: Adapting Video Diffusion Models to World Models
Rigter, M., Gupta, T., Hilmkil, A., and Ma, C · 2024
Later among the works it cites.
Rolling Diffusion Models
Ruhe, D., Heek, J., Salimans, T., and Hoogeboom, E · 2024
Later among the works it cites.
Learning to Act without Actions
Schmidt, D. and Jiang, M · 2024
Later among the works it cites.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
Video Creation by Demonstration
Sun, Y., Zhou, H., Yuan, L., Sun, J. J., Li, Y., Jia, X., Adam, H., Hariharan, B., Zhao, L., and Liu, T · 2024
Later among the works it cites.
DriveDreamer: Towards Real-World-Driven World Models for Autonomous Driving
Wang, X., Zhu, Z., Huang, G., Chen, X., Zhu, J., and Lu, J · 2024
Later among the works it cites.
Jafar: An Open-Source Genie Reimplemention in JAX
Willi, T., Jackson, M. T., and Foerster, J. N · 2024
Later among the works it cites.
iVideoGPT: Interactive VideoGPTs are Scalable World Models
Wu, J., Yin, S., Feng, N., He, X., Li, D., Hao, J., and Long, M · 2024
Later among the works it cites.
Pandora: Towards General World Model with Natural Language Actions and Video States
Xiang, J., Liu, G., Gu, Y., Gao, Q., Ning, Y., Zha, Y., Feng, Z., Tao, T., Hao, S., Shi, Y., et al · 2024
Later among the works it cites.
Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer
Yatim, D., Fridman, R., Bar-Tal, O., Kasten, Y., and Dekel, T · 2024
Later among the works it cites.
PreLAR: World Model Pre-Training with Learnable Action Representation
Zhang, L., Kan, M., Shan, S., and Chen, X · 2024
Later among the works it cites.
3D-VLA: A 3D Vision-Language-Action Generative World Model
Zhen, H., Qiu, X., Chen, P., Yang, J., Yan, X., Du, Y., Hong, Y., and Gan, C · 2024
Later among the works it cites.
IRASim: Learning Interactive Real-Robot Action Simulators
Zhu, F., Wu, H., Guo, S., Liu, Y., Cheang, C., and Kong, T · 2024
Later among the works it cites.
Cosmos World Foundation Model Platform for Physical AI
Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al · 2025
Closest in time.
Navigation World Models
Bar, A., Zhou, G., Tran, D., Darrell, T., and LeCun, Y · 2025
Closest in time.
Bu, Q., Cai, J., Chen, L., Cui, X., Ding, Y., Feng, S., Gao, S., He, X., Huang, X., Jiang, S., et al · 2025
Closest in time.
GameGen-X: Interactive Open-World Game Video Generation
Che, H., He, X., Liu, Q., Jin, C., and Chen, H · 2025
Closest in time.
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Hassan, M., Stapf, S., Rahimi, A., Rezende, P., Haghighi, Y., Brüggemann, D., Katircioglu, I., Zhang, L., Chen, X., Saha, S., et al · 2025
Closest in time.
Pre-Trained Video Generative Models as World Simulators
He, H., Zhang, Y., Lin, L., Xu, Z., and Pan, L · 2025
Closest in time.
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
Hong, Y., Liu, B., Wu, M., Zhai, Y., Chang, K.-W., Li, L., Lin, K., Lin, C.-C., Wang, J., Yang, Z., et al · 2025
Closest in time.
MotionClone: Training-Free Motion Cloning for Controllable Video Generation
Ling, P., Bu, J., Zhang, P., Dong, X., Zang, Y., Wu, T., Chen, H., Wang, J., and Jin, Y · 2025
Closest in time.
Generative World Explorer
Lu, T., Shu, T., Yuille, A., Khashabi, D., and Chen, J · 2025
Closest in time.
Latent Action Learning Requires Supervision in the Presence of Distractors
Nikulin, A., Zisman, I., Tarasov, D., Lyubaykin, N., Polubarov, A., Kiselev, I., and Kurenkov, V · 2025
Closest in time.
Strengthening Generative Robot Policies through Predictive World Modeling
Qi, H., Yin, H., Du, Y., and Yang, H · 2025
Closest in time.
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
Ren, Z., Wei, Y., Guo, X., Zhao, Y., Kang, B., Feng, J., and Jin, X · 2025
Closest in time.
Diffusion Models are Real-Time Game Engines
Valevski, D., Leviathan, Y., Arar, M., and Fruchter, S · 2025
Closest in time.
Villar-Corrales, A. and Behnke, S · 2025
Closest in time.
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
Wang, L., Zhao, K., Liu, C., and Chen, X · 2025
Closest in time.
From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment
Wu, Y., Tian, R., Swamy, G., and Bajcsy, A · 2025
Closest in time.
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al · 2025
Closest in time.
Latent Action Pretraining from Videos
Ye, S., Jang, J., Jeon, B., Joo, S., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.-W., Lin, B. Y., et al · 2025
Closest in time.
From Slow Bidirectional to Fast Causal Video Generators
Yin, T., Zhang, Q., Zhang, R., Freeman, W. T., Durand, F., Shechtman, E., and Huang, X · 2025
Closest in time.
GameFactory: Creating New Games with Generative Interactive Videos
Yu, J., Qin, Y., Wang, X., Wan, P., Zhang, D., and Liu, X · 2025
Closest in time.