Fetching the paper…
Reading the bibliography…
The focus of this work is sample-efficient deep reinforcement learning (RL) with a simulator.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
J. Sherman and W. J. Morrison · 1950
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Learning to act using real-time dynamic programming
A. G. Barto, S. J. Bradtke, and S. P. Singh · 1995
Earlier work this paper cites.
Never give up: Learning directed exploration strategies
A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, et al · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
Bounded real-time dynamic programming: RTDP with monotone upper bounds and performance guarantees
H. B. McMahan, M. Likhachev, and G. J. Gordon · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo tree search
R. Coulom · 2006
Earlier work this paper cites.
Bandit based Monte-Carlo planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Focused real-time dynamic programming for MDPs: Squeezing more out of a heuristic
T. Smith and R. Simmons · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
Bayesian real-time dynamic programming
S. Sanner, R. Goetschalckx, K. Driessens, and G. Shani · 2009
Earlier work this paper cites.
Modeling and simulation of 5 dof educational robot arm
M. A. Qassem, I. Abuhadrous, and H. Elaydi · 2010
Earlier work this paper cites.
Approximate policy iteration: A survey and some new methods
D. P. Bertsekas · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, et al · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2019
Later among the works it cites.
Behaviour suite for reinforcement learning
I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepesvari, S. Singh, et al · 2019
Later among the works it cites.
Exploring restart distributions
A. Tavakoli, V. Levdik, R. Islam, C. M. Smith, and P. Kormushev · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
L. Yang and M. Wang · 2019
Later among the works it cites.
PC-PG: Policy cover directed exploration for provable policy gradient learning
A. Agarwal, M. Henaff, S. Kakade, and W. Sun · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning with double Q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
UCB exploration via Q-ensembles
R. Y. Chen, S. Sidor, P. Abbeel, and J. Schulman · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. Xi Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel · 2017
Cited alongside, same era.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2018
Cited alongside, same era.
Survey of deep reinforcement learning for motion planning of autonomous vehicles
S. Aradi · 2020
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
S. S. Du, S. M. Kakade, R. Wang, and L. F. Yang · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
M. W. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, N. Momchev, D. Sinopalnikov, P. Stańczyk, S. Ramos, A. Raichuk, D. Vincent, L. Hussenot, R. Dadashi, G. Dulac-Arnold, M. Orsini, A. Jacq, J. Ferret, N. Vieillard, S. K. S. Ghasemipour, S. Girgin, O. Pietquin, F. Behbahani, T. Norman, A. Abdolmaleki, A. Cassirer, F. Yang, K. Baumli, S. Henderson, A. Friesen, R. Haroun, A. Novikov, S. G. Colmenarejo, S. Cabi, C. Gulcehre, T. L. Paine, S. Srinivasan, A. Cowie, Z. Wang, B. Piot, and N. de Freitas · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan · 2020
Later among the works it cites.
First return, then explore
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2021
Later among the works it cites.
Improved regret bound and experience replay in regularized policy iteration
N. Lazic, D. Yin, Y. Abbasi-Yadkori, and C. Szepesvari · 2021
Later among the works it cites.
G. Li, Y. Chen, Y. Chi, Y. Gu, and Y. Wei · 2021
Later among the works it cites.
Cautiously optimistic policy optimization and exploration with linear function approximation
A. Zanette, C.-A. Cheng, and A. Agarwal · 2021
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
J. Degrave, F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de Las Casas, et al · 2022
Later among the works it cites.
Confident least square value iteration with local access to a simulator
B. Hao, N. Lazić, D. Yin, Y. Abbasi-Yadkori, and C. Szepesvári · 2022
Later among the works it cites.
S. Kapturowski, V. Campos, R. Jiang, N. Rakićević, H. van Hasselt, C. Blundell, and A. P. Badia · 2022
Later among the works it cites.
G. Weisz, A. György, T. Kozuno, and C. Szepesvári · 2022
Later among the works it cites.
Efficient local planning with linear function approximation
D. Yin, B. Hao, Y. Abbasi-Yadkori, N. Lazić, and C. Szepesvári · 2022
Later among the works it cites.
Can agents run relay race with strangers? generalization of RL to out-of-distribution trajectories
L.-C. Lan, H. Zhang, and C.-J. Hsieh · 2023
Closest in time.