Fetching the paper…
Reading the bibliography…
We introduce Diffusion World Model (DWM), a conditional diffusion model capable of predicting multistep future states and rewards concurrently.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. (2019) · 1903
Earlier work this paper cites.
Combating the compounding-error problem with a multi-step model
Asadi, K., Misra, D., Kim, S., and Littman, M. L. (2019) · 1905
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Tucker, G., and Levine, S. (2019) · 1906
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019) · 1907
Earlier work this paper cites.
Benchmarking model-based reinforcement learning
Wang, T., Bao, X., Clavera, I., Hoang, J., Wen, Y., Langlois, E., Zhang, S., Zhang, G., Abbeel, P., and Ba, J. (2019) · 1907
Earlier work this paper cites.
A self regularized non-monotonic activation function [j]
Mish, M. D. (2019) · 1908
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S. (2019) · 1910
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O. (2019) · 1911
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. (2019a) · 1912
Earlier work this paper cites.
Learning to combat compounding-error in model-based reinforcement learning
Xiao, C., Wu, Y., Ma, C., Schuurmans, D., and Müller, M. (2019) · 1912
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. (1990) · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. (1991) · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A. (2014) · 1993
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2004
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2006
Earlier work this paper cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J. (2020) · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. (2020) · 2011
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., Peters, J., et al. (2013) · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T. (2015) · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2015) · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. (2015) · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Williams, G., Aldrich, A., and Theodorou, E. (2015) · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D. (2018) · 2018
Earlier work this paper cites.
Ha, D. and Schmidhuber, J. (2018) · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S. (2018) · 2018
Cited alongside, same era.
Group normalization
Wu, Y. and He, K. (2018) · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S. (2019) · 2019
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S. (2020) · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P. (2020) · 2020
Flow matching for generative modeling
Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. (2022) · 2022
Later among the works it cites.
Transformers are sample efficient world models
Micheli, V., Alonso, E., and Fleuret, F. (2022) · 2022
Later among the works it cites.
Conserweightive behavioral cloning for reliable offline reinforcement learning
Nguyen, T., Zheng, Q., and Grover, A. (2022) · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M. (2022) · 2022
Later among the works it cites.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
gamma-models: Generative temporal difference learning for infinite-horizon prediction
Janner, M., Mordatch, I., and Levine, S. (2020) · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. (2020) · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020) · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T. (2020) · 2020
Cited alongside, same era.
On the model-based stochastic value gradient for continuous reinforcement learning
Amos, B., Stanton, S., Yarats, D., and Wilson, A. G. (2021) · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Cited alongside, same era.
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S. (2023) · 2023
Later among the works it cites.
Iql-td-mpc: Implicit q-learning for hierarchical model predictive control
Chitnis, R., Xu, Y., Hashemi, B., Lehnert, L., Dogan, U., Zhu, Z., and Delalleau, O. (2023) · 2023
Later among the works it cites.
Consistency models as a rich and efficient policy class for reinforcement learning
Ding, Z. and Jin, C. (2023) · 2023
Later among the works it cites.
Learning universal policies via text-guided video generation
Du, Y., Yang, M., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., and Abbeel, P. (2023) · 2023
Later among the works it cites.
Extreme q-learning: Maxent rl without entropy
Garg, D., Hejna, J., Geist, M., and Ermon, S. (2023) · 2023
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. (2023) · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, N., Su, H., and Wang, X. (2023) · 2023
Later among the works it cites.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P., Kostrikov, I., Janner, M., Kuba, J. G., and Levine, S. (2023) · 2023
Later among the works it cites.
Chain-of-thought predictive control
Jia, Z., Liu, F., Thumuluri, V., Chen, L., Huang, Z., and Su, H. (2023) · 2023
Later among the works it cites.
Lu, C., Ball, P. J., and Parker-Holder, J. (2023) · 2023
Later among the works it cites.
Reorientdiff: Diffusion model based reorientation for object manipulation
Mishra, U. A. and Chen, Y. (2023) · 2023
Later among the works it cites.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Nakamoto, M., Zhai, Y., Singh, A., Mark, M. S., Ma, Y., Finn, C., Kumar, A., and Levine, S. (2023) · 2023
Later among the works it cites.
World models via policy-guided trajectory diffusion
Rigter, M., Yamada, J., and Posner, I. (2023) · 2023
Later among the works it cites.
Transformer-based world models are happy with 100k interactions
Robine, J., Höftmann, M., Uelwer, T., and Harmeling, S. (2023) · 2023
Later among the works it cites.
Song, Y., Dhariwal, P., Chen, M., and Sutskever, I. (2023) · 2023
Later among the works it cites.
Controllable video generation by learning the underlying dynamical system with neural ode
Xu, Y., Li, N., Goel, A., Guo, Z., Yao, Z., Kasaei, H., Kasaei, M., and Li, Z. (2023) · 2023
Later among the works it cites.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Yamagata, T., Khalil, A., and Santos-Rodriguez, R. (2023) · 2023
Later among the works it cites.
Learning interactive real-world simulators
Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P. (2023) · 2023
Later among the works it cites.
Learning unsupervised world models for autonomous driving via discrete diffusion
Zhang, L., Xiong, Y., Yang, Z., Casas, S., Hu, R., and Urtasun, R. (2023) · 2023
Later among the works it cites.
Diffusion world models
Alonso, E., Jelley, A., Kanervisto, A., and Pearce, T. (2024) · 2024
Closest in time.
Jackson, M. T., Matthews, M. T., Lu, C., Ellis, B., Whiteson, S., and Foerster, J. (2024) · 2024
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D. (2019) · 2062
Closest in time.