Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) combines a control problem with statistical estimation: The system dynamics are not known to the agent, but can be learned through experience.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
The advanced theory of statistics
Maurice George Kendall · 1946
Earlier work this paper cites.
Statistical decision functions
Abraham Wald · 1950
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John C Gittins · 1979
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
An introduction to the Kalman filter
Greg Welch, Gary Bishop, et al · 1995
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 2005
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Probabilistic inference for solving discrete and continuous state markov decision processes
Marc Toussaint and Amos Storkey · 2006
Earlier work this paper cites.
Stochastic simulation: Algorithms and analysis , volume 57
Søren Asmussen and Peter W Glynn · 2007
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
General duality between optimal control and estimation
Emanuel Todorov · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan R Peters · 2009
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Efficient computation of optimal actions
Emanuel Todorov · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
Variational methods for reinforcement learning
Thomas Furmston and David Barber · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Cited alongside, same era.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li · 2011
Cited alongside, same era.
Efficient Bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Cited alongside, same era.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Cited alongside, same era.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Cited alongside, same era.
Boltzmann exploration done right
Nicolò Cesa-Bianchi, Claudio Gentile, Gergely Neu, and Gabor Lugosi · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Combining policy gradient and Q-learning
Brendan O’Donoghue, Remi Munos, Koray Kavukcuoglu, and Volodymyr Mnih · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning
Ian Osband and Benjamin Van Roy · 2017
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Daniel Russo, Zheng Wen, and Benjamin Van Roy · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos Nicolas Heess, and Martin Riedmiller · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
From bandits to monte-carlo tree search: The optimistic principle applied to optimization and planning
Rémi Munos · 2014
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2014
Cited alongside, same era.
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Variational Bayesian reinforcement learning with regret bounds
Brendan O’Donoghue · 2018
Later among the works it cites.
The uncertainty Bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Later among the works it cites.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew Maas, J Andrew Bagnell, and Anind K Dey · 2018
Later among the works it cites.
Virel: A variational inference framework for reinforcement learning
Matthew Fellows, Anuj Mahajan, Tim GJ Rudner, and Shimon Whiteson · 2019
Later among the works it cites.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, , Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepezvari, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado Van Hasselt · 2019
Later among the works it cites.