Fetching the paper…
Reading the bibliography…
Research in model-based reinforcement learning has made significant progress in recent years.
Stochastic games
Shapley, L. S. 1953 · 1953
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. 1990 · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. 1991 · 1991
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Boutilier, C. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Model-based reinforcement learning: A survey
Moerland, T. M.; Broekens, J.; and Jonker, C. M. 2020 · 2006
Earlier work this paper cites.
A survey of numerical methods for optimal control
Rao, A. V. 2009 · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Earlier work this paper cites.
Model predictive control
Camacho, E. F.; and Alba, C. B. 2013 · 2013
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P.; Neumann, G.; Peters, J.; et al. 2013 · 2013
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Gheshlaghi Azar, M.; Munos, R.; and Kappen, H. J. 2013 · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Earlier work this paper cites.
A concise introduction to decentralized POMDPs
Oliehoek, F. A.; and Amato, C. 2016 · 2016
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sukhbaatar, S.; Fergus, R.; et al. 2016 · 2016
Earlier work this paper cites.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G.; and Boedecker, J. 2017 · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R.; Wu, Y. I.; Tamar, A.; Harb, J.; Pieter Abbeel, O.; and Mordatch, I. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K.; Calandra, R.; McAllister, R.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V.; Wan, A.; Stoica, I.; Jordan, M. I.; Gonzalez, J. E.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Communication in multi-agent reinforcement learning: Intention sharing
Kim, W.; Park, J.; and Sung, Y. 2020 · 2020
Later among the works it cites.
Multi-agent game abstraction via graph attention neural network
Liu, Y.; Wang, W.; Hu, Y.; Hao, J.; Chen, X.; and Gao, Y. 2020 · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Nagabandi, A.; Konolige, K.; Levine, S.; and Kumar, V. 2020 · 2020
Later among the works it cites.
Trust the model when it is confident: Masked model-based actor-critic
Pan, F.; He, J.; Tu, D.; and He, Q. 2020 · 2020
Later among the works it cites.
From few to more: Large-scale dynamic multiagent curriculum learning
Wang, W.; Yang, T.; Liu, Y.; Hao, J.; Hao, X.; Hu, Y.; Chen, Y.; Fan, C.; and Gao, Y. 2020 · 2020
Later among the works it cites.
Error bounds of imitating policies and environments
Xu, T.; Li, Z.; and Yu, Y. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Counterfactual multi-agent policy gradients
Foerster, J.; Farquhar, G.; Afouras, T.; Nardelli, N.; and Whiteson, S. 2018 · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Kurutach, T.; Clavera, I.; Duan, Y.; Tamar, A.; and Abbeel, P. 2018 · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T.; Samvelyan, M.; Schroeder, C.; Farquhar, G.; Foerster, J.; and Whiteson, S. 2018 · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D.; Lillicrap, T.; Fischer, I.; Villegas, R.; Ha, D.; Lee, H.; and Davidson, J. 2019 · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M.; Fu, J.; Zhang, M.; and Levine, S. 2019 · 2019
Cited alongside, same era.
Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
Luo, Y.; Xu, H.; Li, Y.; Tian, Y.; Darrell, T.; and Ma, T. 2019 · 2019
Cited alongside, same era.
The StarCraft Multi-Agent Challenge
Samvelyan, M.; Rashid, T.; de Witt, C. S.; Farquhar, G.; Nardelli, N.; Rudner, T. G.; Hung, C.-M.; Torr, P. H.; Foerster, J. N.; and Whiteson, S. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
On the model-based stochastic value gradient for continuous reinforcement learning
Amos, B.; Stanton, S.; Yarats, D.; and Wilson, A. G. 2021 · 2021
Later among the works it cites.
Pc-mlp: Model-based reinforcement learning with policy cover guided exploration
Song, Y.; and Sun, W. 2021 · 2021
Later among the works it cites.
QPLEX: Duplex Dueling Multi-Agent Q-Learning
Wang, J.; Ren, Z.; Liu, T.; Yu, Y.; and Zhang, C. 2021 · 2021
Later among the works it cites.
Coordinated proximal policy optimization
Wu, Z.; Yu, C.; Ye, D.; Zhang, J.; Zhuo, H. H.; et al. 2021 · 2021
Later among the works it cites.
Model-based multi-agent policy optimization with adaptive opponent-wise rollouts
Zhang, W.; Wang, X.; Shen, J.; and Zhou, M. 2021 · 2021
Later among the works it cites.
Adversarial Counterfactual Environment Model Learning
Chen, X.-H.; Yu, Y.; Zhu, Z.-M.; Yu, Z.; Chen, Z.; Wang, C.; Wu, Y.; Wu, H.; Qin, R.-J.; Ding, R.; et al. 2022 · 2022
Later among the works it cites.
Scalable Multi-Agent Model-Based Reinforcement Learning
Egorov, V.; and Shpilman, A. 2022 · 2022
Later among the works it cites.
Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning
Fu, W.; Yu, C.; Xu, Z.; Yang, J.; and Wu, Y. 2022 · 2022
Later among the works it cites.
A Survey on Model-based Reinforcement Learning
Luo, F.-M.; Xu, T.; Lai, H.; Chen, X.-H.; Zhang, W.; and Yu, Y. 2022 · 2022
Later among the works it cites.
Model-based Multi-agent Reinforcement Learning: Recent Progress and Prospects
Wang, X.; Zhang, Z.; and Zhang, W. 2022 · 2022
Later among the works it cites.