Fetching the paper…
Reading the bibliography…
World models constitute a promising approach for training reinforcement learning agents in a safe and sample-efficient manner.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. (2019) · 1903
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Reverse-time diffusion equation models
Anderson, B. D. (1982) · 1982
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D. (1989) · 1989
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. (1991) · 1991
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Computer methods for ordinary differential equations and differential-algebraic equations
Ascher, U. M. and Petzold, L. R. (1998) · 1998
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Gers, F. A., Schmidhuber, J., and Cummins, F. (2000) · 2000
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. (2005) · 2005
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. (2008) · 2008
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P. (2011) · 2011
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Earlier work this paper cites.
Kaiser, Ł. and Sutskever, I. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T. (2015) · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. (2015) · 2015
Earlier work this paper cites.
3d u-net: learning dense volumetric segmentation from sparse annotation
Çiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., and Ronneberger, O. (2016) · 2016
Earlier work this paper cites.
Santana, E. and Hotz, G. (2016) · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. (2016) · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. (2017) · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K. (2018) · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J. (2018) · 2018
Earlier work this paper cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. (2018) · 2018
Earlier work this paper cites.
Group normalization
Wu, Y. and He, K. (2018) · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. (2018) · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Cited alongside, same era.
Neural game engine: Accurate learning of generalizable forward models from pixels
Bamford, C. and Lucas, S. M. (2020) · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. (2020) · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P. (2020) · 2020
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Later among the works it cites.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M. (2022) · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M. (2022) · 2022
Later among the works it cites.
Diffusion world models
Alonso, E., Jelley, A., Kanervisto, A., and Pearce, T. (2023) · 2023
Later among the works it cites.
Muse: Text-to-image generation via masked generative transformers
Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K. P., Freeman, W. T., Rubinstein, M., Li, Y., and Krishnan, D. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to simulate dynamic environments with gamegan
Kim, S. W., Zhou, Y., Philion, J., Torralba, A., and Fidler, S. (2020) · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020) · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. (2020) · 2020
Cited alongside, same era.
Learning semantic-aware normalization for generative adversarial networks
Zheng, H., Fu, J., Zeng, Y., Luo, J., and Zha, Z.-J. (2020) · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M. (2021) · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A. (2021) · 2021
Cited alongside, same era.
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. (2023) · 2023
Later among the works it cites.
Gaia-1: A generative world model for autonomous driving
Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G. (2023) · 2023
Later among the works it cites.
Adaptdiffuser: Diffusion models as adaptive self-evolving planners
Liang, Z., Mu, Y., Ding, M., Ni, F., Tomizuka, M., and Luo, P. (2023) · 2023
Later among the works it cites.
Lu, C., Ball, P. J., and Parker-Holder, J. (2023) · 2023
Later among the works it cites.
Diffusion hyperfeatures: Searching through time and space for semantic correspondence
Luo, G., Dunlap, L., Park, D. H., Holynski, A., and Darrell, T. (2023) · 2023
Later among the works it cites.
Transformers are sample-efficient world models
Micheli, V., Alonso, E., and Fleuret, F. (2023) · 2023
Later among the works it cites.
Extracting reward functions from diffusion models
Nuti, F., Franzmeyer, T., and Henriques, J. F. (2023) · 2023
Later among the works it cites.
Imitating human behaviour with diffusion models
Pearce, T., Rashid, T., Kanervisto, A., Bignell, D., Sun, M., Georgescu, R., Macua, S. V., Tan, S. Z., Momennejad, I., Hofmann, K., and Devlin, S. (2023) · 2023
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S. (2023) · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R. (2023) · 2023
Later among the works it cites.
Transformer-based world models are happy with 100k interactions
Robine, J., Höftmann, M., Uelwer, T., and Harmeling, S. (2023) · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S. (2023) · 2023
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al. (2023) · 2023
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual descriptions
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D. (2023) · 2023
Later among the works it cites.
Daydreamer: World models for physical robot learning
Wu, P., Escontrela, A., Hafner, D., Abbeel, P., and Goldberg, K. (2023) · 2023
Later among the works it cites.
Open-vocabulary panoptic segmentation with text-to-image diffusion models
Xu, J., Liu, S., Vahdat, A., Byeon, W., Wang, X., and De Mello, S. (2023) · 2023
Later among the works it cites.
Temporally consistent transformers for video generation
Yan, W., Hafner, D., James, S., and Abbeel, P. (2023) · 2023
Later among the works it cites.
Storm: Efficient stochastic transformer based world models for reinforcement learning
Zhang, W., Wang, G., Sun, J., Yuan, Y., and Huang, G. (2023) · 2023
Later among the works it cites.
Lumiere: A space-time diffusion model for video generation
Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Li, Y., Michaeli, T., et al. (2024) · 2024
Closest in time.
Video generation models as world simulators
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A. (2024) · 2024
Closest in time.
Genie: Generative interactive environments
Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al. (2024) · 2024
Closest in time.
Ding, Z., Zhang, A., Tian, Y., and Zheng, Q. (2024) · 2024
Closest in time.
Jackson, M. T., Matthews, M. T., Lu, C., Ellis, B., Whiteson, S., and Foerster, J. (2024) · 2024
Closest in time.
Diffusion models are real-time game engines
Valevski, D., Leviathan, Y., Arar, M., and Fruchter, S. (2024) · 2024
Closest in time.