Fetching the paper…
Reading the bibliography…
Recently deep reinforcement learning (DRL) has achieved outstanding success on solving many difficult and large-scale RL problems.
Tsallis’ entropy, ehrenfest theorem and information theory
A. R. Plastino and A. Plastino · 1993
Earlier work this paper cites.
Nonextensive physics: a possible connection between generalized statistical mechanics and quantum groups
C. Tsallis · 1994
Earlier work this paper cites.
Dynamic programming and optimal control
D. P. Bertsekas · 1995
Earlier work this paper cites.
Asymptopia: an exposition of statistical asymptotic theory. 2000
D. Pollard · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Incorporating second-order functional knowledge for better option pricing
C. Dugas, Y. Bengio, F. Bélisle, C. Nadeau, and R. Garcia · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. Kakade · 2003
Earlier work this paper cites.
Pac model-free reinforcement learning
A. L. Strehl, L. Li, E. Wiewiora, J. Langford, and M. L. Littman · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
P. Auer and R. Ortner · 2007
Earlier work this paper cites.
Probably approximately correct (PAC) exploration in reinforcement learning
A. L. Strehl · 2007
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
D. Russo and B. Van Roy · 2013
Earlier work this paper cites.
Generalization and exploration via randomized value functions
I. Osband, B. Van Roy, and Z. Wen · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2014
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2015
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
C. Dann and E. Brunskill · 2015
Cited alongside, same era.
Bayesian reinforcement learning: A survey
M. Ghavamzadeh, S. Mannor, J. Pineau, A. Tamar, et al · 2015
Cited alongside, same era.
UCB Exploration via Q-Ensembles
R. Y. Chen, S. Sidor, P. Abbeel, and J. Schulman · 2017
Later among the works it cites.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
K. Lee, S. Choi, and S. Oh · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjelandand G. Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Deep Reinforcement Learning with Double Q-Learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Traffic signal timing via deep reinforcement learning
L. Li, Y. Lv, and F. Y. Wang · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning
I. Osband and B. Van Roy · 2016
Cited alongside, same era.
Combining policy gradient and Q-learning
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih · 2017
Later among the works it cites.
Deep exploration via randomized value functions
I. Osband, D. Russo, Z. Wen, and B. Van Roy · 2017
Later among the works it cites.
Equivalence between policy gradients and soft Q-Learning
J. Schulman, P. Abbeel, and X. Chen · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Y. Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba · 2017
Later among the works it cites.
An adaptive clipping approach for proximal policy optimization
G. Chen, Y. Peng, and M. Zhang · 2018
Closest in time.
Constrained expectation-maximization methods for effective reinforcement learning
G. Chen, Y. Peng, and M. Zhang · 2018
Closest in time.
Path Consistency Learning in Tsallis Entropy Regularized MDPs
O. Nachum, Y. Chow, and M. Ghavamzadeh · 2018
Closest in time.