Fetching the paper…
Reading the bibliography…
We study the optimal sample complexity in large-scale Reinforcement Learning (RL) problems with policy space generalization, i.e.
Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
N. Littlestone · 1988
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2002
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
A. Tewari and P. L. Bartlett · 2008
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
D. Russo and B. Van Roy · 2013
Earlier work this paper cites.
Analysis of boolean functions
R. O’Donnell · 2014
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
I. Osband and B. Van Roy · 2014
Earlier work this paper cites.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Cited alongside, same era.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, et al · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2017
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
I. Osband and B. Van Roy · 2017
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2019
Later among the works it cites.
Private PAC learning implies finite Littlestone dimension
N. Alon, R. Livni, M. Malliaris, and S. Moran · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
J. Bhandari and D. Russo · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Q. Cai, Z. Yang, C. Jin, and Z. Wang · 2019
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
S. S. Du, S. M. Kakade, R. Wang, and L. F. Yang · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Osband, B. Van Roy, D. Russo, and Z. Wen · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Efficient reinforcement learning in deterministic systems with value function generalization
Z. Wen and B. Van Roy · 2017
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
M. Fazel, R. Ge, S. M. Kakade, and M. Mesbahi · 2018
Cited alongside, same era.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
D. Malik, A. Pananjady, K. Bhatia, K. Khamaru, P. L. Bartlett, and M. J. Wainwright · 2018
Cited alongside, same era.
Information-directed exploration for deep reinforcement learning
N. Nikolov, J. Kirschner, F. Berkenkamp, and A. Krause · 2018
Cited alongside, same era.
W. Sun, N. Jiang, A. Krishnamurthy, A. Agarwal, and J. Langford · 2018
Cited alongside, same era.
Later among the works it cites.
Provably efficient Q-learning with function approximation via distribution shift error checking oracle
S. S. Du, Y. Luo, R. Wang, and H. Zhang · 2019
Later among the works it cites.
Optimistic policy optimization via multiple importance sampling
M. Papini, A. M. Metelli, L. Lupo, and M. Restelli · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
D. Russo · 2019
Later among the works it cites.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
J. Schmidhuber · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
M. Simchowitz and K. G. Jamieson · 2019
Later among the works it cites.
Transfer of samples in policy search via multiple importance sampling
A. Tirinzoni, M. Salvini, and M. Restelli · 2019
Later among the works it cites.
S. S. Du, J. D. Lee, G. Mahajan, and R. Wang · 2020
Closest in time.