Fetching the paper…
Reading the bibliography…
We propose a novel policy update that combines regularized policy optimization with model learning as an auxiliary loss.
Optimal Control of Markov Processes with Incomplete State Information I
Åström, K · 1965
Earlier work this paper cites.
Paper: Model predictive heuristic control
Richalet, J., Rault, A., Testud, J. L., and Papon, J · 1978
Earlier work this paper cites.
Learning how the world works: Specifications for predictive networks in robots and brains
Werbos, P. J · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Schmidhuber, J · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning and reacting based on dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. and Peng, J · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R · 1992
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable Markov decision problems
Jaakkola, T., Singh, S. P., and Jordan, M. I · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
Singh, S. P., Jaakkola, T., and Jordan, M. I · 1994
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Tesauro, G · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Gnugo, 2005
Bump et al · 2005
Earlier work this paper cites.
Neural fitted Q iteration – first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search
Coulom, R · 2006
Earlier work this paper cites.
Reinforcement learning in continuous action spaces
van Hasselt, H. and Wiering, M. A · 2007
Earlier work this paper cites.
Hierarchical task and motion planning in the now
Kaelbling, L. P. and Lozano-Pérez, T · 2010
Earlier work this paper cites.
Monte-Carlo Planning in Large POMDPs
Silver, D. and Veness, J · 2010
Earlier work this paper cites.
Pachi: State of the art open source Go program
Baudiš, P. and Gailly, J.-l · 2011
Earlier work this paper cites.
On the role of planning in model-based deep reinforcement learning
Hamrick, J. B., Friesen, A. L., Behbahani, F., Guez, A., Viola, F., Witherspoon, S., Anthony, T., Buesing, L., Veličković, P., and Weber, T · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Scaling-up Knowledge for a Cognizant Robot
Degris, T. and Modayil, J · 2012
Earlier work this paper cites.
Model-free reinforcement learning with continuous action in practice
Degris, T., Pilarski, P. M., and Sutton, R. S · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in Atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Learning to Predict Independent of Span
van Hasselt, H. and Sutton, R. S · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H., and Silver, D · 2018
Later among the works it cites.
On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2019
Later among the works it cites.
Shaping belief states with generative environment models for rl
Gregor, K., Jimenez Rezende, D., Besse, F., Wu, Y., Merzic, H., and van den Oord, A · 2019
Later among the works it cites.
An investigation of model-free planning
Guez, A., Mirza, M., Gregor, K., Kabra, R., Racanière, S., Weber, T., Raposo, D., Santoro, A., Orseau, L., Eccles, T., et al · 2019
Later among the works it cites.
Analogues of mental simulation and imagination in deep learning
Hamrick, J. B · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Cited alongside, same era.
Value iteration networks
Tamar, A., WU, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Learning values across many orders of magnitude
van Hasselt, H. P., Guez, A., Guez, A., Hessel, M., Mnih, V., and Silver, D · 2016
Cited alongside, same era.
Sample Efficient Actor-Critic with Experience Replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Hessel, M., van Hasselt, H., Modayil, J., and Silver, D · 2019
Later among the works it cites.
When to Trust Your Model: Model-Based Policy Optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Later among the works it cites.
Model-Based Reinforcement Learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H · 2019
Later among the works it cites.
OpenSpiel: A Framework for Reinforcement Learning in Games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., Hennes, D., Morrill, D., Muller, P., Ewalds, T., Faulkner, R., Kramár, J., De Vylder, B., Saeta, B., Bradbury, J., Ding, D., Borgeaud, S., Lai, M., Schrittwieser, J., Anthony, T., Hughes, E., Danihelka, I., and Ryan-Davis, J · 2019
Later among the works it cites.
Normalizing Flows for Probabilistic Modeling and Inference
Papamakarios, G., Nalisnick, E., Jimenez Rezende, D., Mohamed, S., and Lakshminarayanan, B · 2019
Later among the works it cites.
General non-linear Bellman equations
van Hasselt, H., Quan, J., Hessel, M., Xu, Z., Borsa, D., and Barreto, A · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
van Hasselt, H. P., Hessel, M., and Aslanides, J · 2019
Later among the works it cites.
ACE: An actor ensemble algorithm for continuous control with tree search
Zhang, S. and Yao, H · 2019
Later among the works it cites.
Reinforcement Learning: Theory and Algorithms , 2020
Agarwal, A., Jiang, N., Kakade, S. M., and Sun, W · 2020
Later among the works it cites.
Imagined value gradients: Model-based policy optimization with tranferable latent dynamics models
Byravan, A., Springenberg, J. T., Abdolmaleki, A., Hafner, R., Neunert, M., Lampe, T., Siegel, N., Heess, N., and Riedmiller, M · 2020
Later among the works it cites.
Cobbe, K., Hilton, J., Klimov, O., and Schulman, J · 2020
Later among the works it cites.
MCTS as regularized policy optimization
Grill, J.-B., Altché, F., Tang, Y., Hubert, T., Valko, M., Antonoglou, I., and Munos, R · 2020
Later among the works it cites.
The value equivalence principle for model-based reinforcement learning
Grimm, C., Barreto, A., Singh, S., and Silver, D · 2020
Later among the works it cites.
Value-driven hindsight modelling
Guez, A., Viola, F., Weber, T., Buesing, L., Kapturowski, S., Precup, D., Silver, D., and Heess, N · 2020
Later among the works it cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Guo, Z. D., Pires, B. A., Piot, B., Grill, J.-B., Altché, F., Munos, R., and Azar, M. G · 2020
Later among the works it cites.
Mastering Atari with Discrete World Models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Later among the works it cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Model-based Reinforcement Learning: A Survey
Moerland, T. M., Broekens, J., and Jonker, C. M · 2020
Later among the works it cites.
Causally Correct Partial Models for Reinforcement Learning
Rezende, D. J., Danihelka, I., Papamakarios, G., Ke, N. R., Jiang, R., Weber, T., Gregor, K., Merzic, H., Viola, F., Wang, J., Mitrovic, J., Besse, F., Antonoglou, I., and Buesing, L · 2020
Later among the works it cites.
Off-Policy Actor-Critic with Shared Experience Replay
Schmitt, S., Hessel, M., and Simonyan, K · 2020
Later among the works it cites.
Mastering Atari, Go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 2020
Later among the works it cites.
Local Search for Policy Iteration in Continuous Control
Springenberg, J. T., Heess, N., Mankowitz, D., Merel, J., Byravan, A., Abdolmaleki, A., Kay, J., Degrave, J., Schrittwieser, J., Tassa, Y., Buchli, J., Belov, D., and Riedmiller, M · 2020
Later among the works it cites.
Mirror Descent Policy Optimization
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M · 2020
Later among the works it cites.
Leverage the Average: an Analysis of KL Regularization in RL
Vieillard, N., Kozuno, T., Scherrer, B., Pietquin, O., Munos, R., and Geist, M · 2020
Later among the works it cites.
A self-tuning actor-critic algorithm, 2020
Zahavy, T., Xu, Z., Veeriah, V., Hessel, M., Oh, J., van Hasselt, H., Silver, D., and Singh, S · 2020
Later among the works it cites.
Podracer architectures for scalable reinforcement learning
Hessel, M., Kroiss, M., Clark, A., Kemaev, I., Quan, J., Keck, T., Viola, F., and van Hasselt, H · 2021
Closest in time.
Learning and Planning in Complex Action Spaces
Hubert, T., Schrittwieser, J., Antonoglou, I., Barekatain, M., Schmitt, S., and Silver, D · 2021
Closest in time.
Online and offline reinforcement learning by planning with a learned model
Schrittwieser, J., Hubert, T., Mandhane, A., Barekatain, M., Antonoglou, I., and Silver, D · 2021
Closest in time.