Fetching the paper…
Reading the bibliography…
State-of-the-art reinforcement learning algorithms mostly rely on being allowed to directly interact with their environment to collect millions of observations.
Striving for simplicity in off-policy deep reinforcement learning
Agarwal, R.; Schuurmans, D.; and Norouzi, M. 2019 · 1907
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N.; Ghandeharioun, A.; Shen, J. H.; Ferguson, C.; Lapedriza, A.; Jones, N.; Gu, S.; and Picard, R. 2019 · 1907
Earlier work this paper cites.
Benchmarking Batch Deep Reinforcement Learning Algorithms
Fujimoto, S.; Conti, E.; Ghavamzadeh, M.; and Pineau, J. 2019 · 1910
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. 1990 · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. 1992 · 1992
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S.; and Barto, A. G. 1998 · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y.; Russell, S. J.; et al. 2000 · 2000
Earlier work this paper cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y.; Springenberg, J. T.; Berkenkamp, F.; Abdolmaleki, A.; Neunert, M.; Lampe, T.; Hafner, R.; and Riedmiller, M. 2020 · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G.; and Parr, R. 2003 · 2003
Earlier work this paper cites.
Tree-Based Batch Mode Reinforcement Learning
Ernst, D.; Geurts, P.; and Wehenkel, L. 2005 · 2005
Earlier work this paper cites.
Approximate Value Iteration in the Reinforcement Learning Context. Application to Electrical Power System Control
Ernst, D.; Glavic, M.; Geurts, P.; and Wehenkel, L. 2005 · 2005
Earlier work this paper cites.
MOReL: Model-Based Offline Reinforcement Learning
Kidambi, R.; Rajeswaran, A.; Netrapalli, P.; and Joachims, T. 2020 · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. 2005 · 2005
Earlier work this paper cites.
MOPO: Model-based Offline Policy Optimization
Yu, T.; Thomas, G.; Yu, L.; Ermon, S.; Zou, J.; Levine, S.; Finn, C.; and Ma, T. 2020 · 2005
Earlier work this paper cites.
Developmental robotics, optimal artificial curiosity, creativity, music, and the fine arts
Schmidhuber, J. 2006 · 2006
Earlier work this paper cites.
Batch reinforcement learning in a complex domain
Kalyanakrishnan, S.; and Stone, P. 2007 · 2007
Earlier work this paper cites.
Improving Optimality of Neural Rewards Regression for Data-Efficient Batch Near-Optimal Policy Identification
Schneegaß, D.; Udluft, S.; and Martinetz, T. 2007 · 2007
Earlier work this paper cites.
Reinforcement learning for robot soccer
Riedmiller, M.; Gabel, T.; Hafner, R.; and Lange, S. 2009 · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S.; and Bagnell, D. 2010 · 2010
Earlier work this paper cites.
Experience replay for real-time reinforcement learning control
Adam, S.; Busoniu, L.; and Babuska, R. 2011 · 2011
Cited alongside, same era.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M.; and Rasmussen, C. E. 2011 · 2011
Cited alongside, same era.
Agent self-assessment: Determining policy quality without execution
Hans, A.; Duell, S.; and Udluft, S. 2011 · 2011
Cited alongside, same era.
A kernel two-sample test
Gretton, A.; Borgwardt, K. M.; Rasch, M. J.; Schölkopf, B.; and Smola, A. 2012 · 2012
Cited alongside, same era.
Batch reinforcement learning
Lange, S.; Gabel, T.; and Riedmiller, M. 2012 · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Abbeel, O. P.; and Zaremba, W. 2017 · 2017
Later among the works it cites.
Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning
Depeweg, S.; Hernández-Lobato, J. M.; Doshi-Velez, F.; and Udluft, S. 2017 · 2017
Later among the works it cites.
A benchmark environment motivated by industrial control problems
Hein, D.; Depeweg, S.; Tokic, M.; Udluft, S.; Hentschel, A.; Runkler, T. A.; and Sterzing, V. 2017 · 2017
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Laroche, R.; Trichelair, P.; and Combes, R. T. d. 2017 · 2017
Later among the works it cites.
Dart: Noise injection for robust imitation learning
Laskey, M.; Lee, J.; Fox, R.; Dragan, A.; and Goldberg, K. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P.; and Welling, M. 2013 · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D.; Lever, G.; Heess, N.; Degris, T.; Wierstra, D.; and Riedmiller, M. 2014 · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Later among the works it cites.
Implicit quantile networks for distributional reinforcement learning
Dabney, W.; Ostrovski, G.; Silver, D.; and Munos, R. 2018 · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S.; Meger, D.; and Precup, D. 2018 · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Later among the works it cites.
Interpretable policies for reinforcement learning by genetic programming
Hein, D.; Udluft, S.; and Runkler, T. A. 2018 · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, T.; Clavera, I.; Duan, Y.; Tamar, A.; and Abbeel, P. 2018 · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A.; Kahn, G.; Fearing, R. S.; and Levine, S. 2018 · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A.; McGrew, B.; Andrychowicz, M.; Zaremba, W.; and Abbeel, P. 2018 · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
Pong, V.; Gu, S.; Dalal, M.; and Levine, S. 2018 · 2018
Later among the works it cites.
Exploring the limitations of behavior cloning for autonomous driving
Codevilla, F.; Santana, E.; López, A. M.; and Gaidon, A. 2019 · 2019
Later among the works it cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Kumar, A.; Fu, J.; Soh, M.; Tucker, G.; and Levine, S. 2019 · 2019
Later among the works it cites.
Deep Exploration via Randomized Value Functions
Osband, I.; Van Roy, B.; Russo, D. J.; and Wen, Z. 2019 · 2019
Later among the works it cites.
Bayesian decomposition of multi-modal dynamical systems for reinforcement learning
Kaiser, M.; Otte, C.; Runkler, T. A.; and Ek, C. H. 2020 · 2020
Closest in time.
Batch Reinforcement Learning with Hyperparameter Gradients
Lee, B.-J.; Lee, J.; Vrancx, P.; Kim, D.; and Kim, K.-E. 2020 · 2020
Closest in time.