Fetching the paper…
Reading the bibliography…
We consider model-based reinforcement learning (MBRL) in 2-agent, high-fidelity continuous control problems -- an important domain for robots interacting with other agents in the same workspace.
Neural networks for control systems—a survey
K. J. Hunt, D. Sbarbaro, R. Żbikowski, and P. J. Gawthrop · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Layered learning in multiagent systems: A winning approach to robotic soccer
P. Stone · 2000
Earlier work this paper cites.
Value-function reinforcement learning in markov games
M. L. Littman · 2001
Earlier work this paper cites.
Learning first-order markov models for control
P. Abbeel and A. Y. Ng · 2005
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Elements of information theory
T. M. Cover and J. A. Thomas · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Generating long-term trajectories using deep hierarchical networks
S. Zheng, Y. Yue, and J. Hobbs · 2016
Cited alongside, same era.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Desire: Distant future prediction in dynamic scenes with interacting agents
N. Lee, W. Choi, P. Vernaza, C. B. Choy, P. H. Torr, and M. Chandraker · 2017
Cited alongside, same era.
Socially aware motion planning with deep reinforcement learning
Y. F. Chen, M. Everett, M. Liu, and J. P. How · 2017
Cited alongside, same era.
Generative modeling of multimodal multi-human behavior
B. Ivanovic, E. Schmerling, K. Leung, and M. Pavone · 2018
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
I. Mordatch and P. Abbeel · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2018
Later among the works it cites.
Deep imitative models for flexible inference, planning, and control
N. Rhinehart, R. McAllister, and S. Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Schmerling, K. Leung, W. Vollprecht, and M. Pavone · 2017
Cited alongside, same era.
Prediction and control with temporal segment models
N. Mishra, P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Information theoretic MPC for model-based reinforcement learning
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou · 2017
Cited alongside, same era.
Learning and policy search in stochastic dynamical systems with bayesian neural networks
S. Depeweg, J. M. Hernández-Lobato, F. Doshi-Velez, and S. Udluft · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Later among the works it cites.
Learning latent subspaces in variational autoencoders
J. Klys, J. Snell, and R. Zemel · 2018
Later among the works it cites.
Unity: A general platform for intelligent agents
A. Juliani, V.-P. Berges, E. Vckay, Y. Gao, H. Henry, M. Mattar, and D. Lange · 2018
Later among the works it cites.
Modeling the long term future in model-based reinforcement learning
N. R. Ke, A. Singh, A. Touati, A. Goyal, Y. Bengio, D. Parikh, and D. Batra · 2019
Closest in time.