Fetching the paper…
Reading the bibliography…
Many model-based reinforcement learning (RL) methods follow a similar template: fit a model to previously observed data, and then use data from that model for RL or planning.
Combating the compounding-error problem with a multi-step model
Asadi, K., Misra, D., Kim, S., and Littman, M. L. (2019) · 1905
Earlier work this paper cites.
Solving uncertain Markov decision processes
Bagnell, J. A., Ng, A. Y., and Schneider, J. G. (2001) · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Mihatsch, O. and Neuneier, R. (2002) · 2002
Earlier work this paper cites.
Robustness in Markov decision problems with uncertain transition matrices
Nilim, A. and El Ghaoui, L. (2003) · 2003
Earlier work this paper cites.
Planning and execution using inaccurate models with provable guarantees
Vemula, A., Oza, Y., Bagnell, J. A., and Likhachev, M. (2020) · 2003
Earlier work this paper cites.
Morel : Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. (2020) · 2005
Earlier work this paper cites.
MOPO: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T. (2020b) · 2005
Earlier work this paper cites.
Variational model-based policy optimization
Chow, Y., Cui, B., Ryu, M., and Ghavamzadeh, M. (2020) · 2006
Earlier work this paper cites.
Internal rewards mitigate agent boundedness
Sorg, J., Singh, S. P., and Lewis, R. L. (2010) · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D. (2010) · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M. P. and Rasmussen, C. E. (2011) · 2011
Earlier work this paper cites.
The value equivalence principle for model-based reinforcement learning
Grimm, C., Barreto, A., Singh, S., and Silver, D. (2020) · 2011
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Ross, S., Gordon, G. J., and Bagnell, J. A. (2011) · 2011
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Ross, S. and Bagnell, D. (2012) · 2012
Earlier work this paper cites.
Reinforcement learning with misspecified model classes
Joseph, J., Geramifard, A., Roberts, J. W., How, J. P., and Roy, N. (2013) · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Talvitie, E. (2014) · 2014
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Donahue, J., Krähenbühl, P., and Darrell, T. (2016) · 2016
Cited alongside, same era.
Adversarially learned inference
Dumoulin, V., Belghazi, I., Poole, B., Mastropietro, O., Lamb, A., Arjovsky, M., and Courville, A. (2016) · 2016
Learning plannable representations with causal InfoGAN
Kurutach, T., Tamar, A., Yang, G., Russell, S. J., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Universal planning networks: Learning generalizable representations for visuomotor control
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C. (2018) · 2018
Later among the works it cites.
A model-based reinforcement learning with adversarial training for online recommendation
Bai, X., Guan, J., and Wang, H. (2019) · 2019
Later among the works it cites.
TF-Agents: A library for reinforcement learning in tensorflow
Guadarrama, S., Korattikara, A., Ramirez, O., Castro, P., Holly, E., Fishman, S., Wang, K., Gonina, E., Wu, N., Kokiopoulou, E., Sbaiz, L., Smith, J., Bartók, G., Berent, J., Harris, C., Vanhoucke, V., and Brevdo, E. (2018) · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improved techniques for training GANs
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. (2016) · 2016
Cited alongside, same era.
Improved learning of dynamics models for control
Venkatraman, A., Capobianco, R., Pinto, L., Hebert, M., Nardi, D., and Bagnell, J. A. (2016) · 2016
Cited alongside, same era.
Value-aware loss function for model-based reinforcement learning
Farahmand, A.-m., Barreto, A., and Nikovski, D. (2017) · 2017
Cited alongside, same era.
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Cited alongside, same era.
Path integral networks: End-to-end differentiable optimal control
Okada, M., Rigazio, L., and Aoshima, T. (2017) · 2017
Cited alongside, same era.
Information Theoretic MPC for Model-Based Reinforcement Learning
Williams, G., Wagener, N., Goldfain, B., Drews, P., Rehg, J. M., Boots, B., and Theodorou, E. A. (2017) · 2017
Cited alongside, same era.
Differentiable MPC for end-to-end planning and control
Amos, B., Rodriguez, I., Sacks, J., Boots, B., and Kolter, Z. (2018) · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S. (2019) · 2019
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T. (2019) · 2019
Later among the works it cites.
ROBEL: Robotics benchmarks for learning with low-cost robots
Ahn, M., Zhu, H., Hartikainen, K., Ponte, H., Gupta, A., Levine, S., and Kumar, V. (2020) · 2020
Later among the works it cites.
Gan-based planning model in deep reinforcement learning
Chen, S., Jiang, J., Zhang, X., Wu, J., and Lu, G. (2020) · 2020
Later among the works it cites.
Gradient-aware model-based policy search
D’Oro, P., Metelli, A. M., Tirinzoni, A., Papini, M., and Restelli, M. (2020) · 2020
Later among the works it cites.
Objective mismatch in model-based reinforcement learning
Lambert, N., Amos, B., Yadan, O., and Calandra, R. (2020) · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
Rajeswaran, A., Mordatch, I., and Kumar, V. (2020) · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020) · 2020
Later among the works it cites.
Control-oriented model-based reinforcement learning with implicit differentiation
Nikishin, E., Abachi, R., Agarwal, R., and Bacon, P.-L. (2021) · 2021
Closest in time.
Outcome-driven reinforcement learning via variational inference
Rudner, T. G., Pong, V. H., McAllister, R., Gal, Y., and Levine, S. (2021) · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C. (2021) · 2021
Closest in time.