Efficient exploration via state marginal matching
Original
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 1906
Earlier work this paper cites.
Network randomization: A simple technique for generalization in deep reinforcement learning
Original
Kimin Lee, Kibok Lee, Jinwoo Shin, and Honglak Lee · 1910
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Original
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 1912
Earlier work this paper cites.
Markov decision processes
Martin L Puterman · 1990
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
R-max: A general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
All else being equal be empowered
Alexander S Klyubin, Daniel Polani, and Chrystopher L Nehaniv · 2005
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank MDPs
Original
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Provably efficient reinforcement learning for discounted MDPs with feature mapping
Original
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2006
Earlier work this paper cites.
PC-PG: Policy cover directed exploration for provable policy gradient learning
Original
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2007
Earlier work this paper cites.
Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Original
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
PAC bounds for discounted MDPs
Tor Lattimore and Marcus Hutter · 2012
Earlier work this paper cites.
BeBold: Exploration beyond the boundary of explored regions
Original
Tianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu, Kurt Keutzer, Joseph E Gonzalez, and Yuandong Tian · 2012
Earlier work this paper cites.
Nearly minimax optimal reinforcement learning for linear mixture Markov decision processes
Original
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2012
Earlier work this paper cites.
Empowerment: An introduction
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Earlier work this paper cites.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Original
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Openai gym
Original
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Variational intrinsic control
Original
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Is the Bellman residual a bad proxy?
Matthieu Geist, Bilal Piot, and Olivier Pietquin · 2017
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Original
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.