Fetching the paper…
Reading the bibliography…
Balancing exploration and exploitation remains a key challenge in reinforcement learning (RL).
Über eine aufgabe der wahrscheinlichkeitsrechnung betreffend die irrfahrt im straßennetz
G. Pólya · 1921
Earlier work this paper cites.
On brownian motions in n-space
S. Kakutani · 1944
Earlier work this paper cites.
Some problems on random walk in space
A. Dvoretzky and P. Erdős · 1951
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Tutorial on large deviations for the binomial distribution
R. Arratia and L. Gordon · 1989
Earlier work this paper cites.
Probabilistic roadmaps for path planning in high-dimensional configuration spaces
L. E. Kavraki, P. Svestka, J. C. Latombe, and M. H. Overmars · 1996
Earlier work this paper cites.
Path planning in expansive configuration spaces
D. Hsu, J.C. Latombe, and R. Motwani · 1997
Earlier work this paper cites.
Rapidly-exploring random trees: A new tool for path planning
S. M. Lavalle · 1998
Earlier work this paper cites.
RRT-connect: An efficient approach to single-query path planning
J. J. Kuffner and S. M. LaValle · 2000
Earlier work this paper cites.
Randomized kinodynamic planning
S. M. LaValle and J. J. Kuffner · 2001
Earlier work this paper cites.
Randomized kinodynamic motion planning with moving obstacles
D. Hsu, R. Kindel, J.C. Latombe, and S. Rock · 2002
Earlier work this paper cites.
Approaches for heuristically biasing RRT growth
C. Urmson and R. Simmons · 2003
Earlier work this paper cites.
How can we define intrinsic motivation?
P. Y. Oudeyer and F. Kaplan · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
The many faces of optimism: a unifying approach
I. Szita and A. Lőrincz · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L.N. Bottou · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
Sampling-based algorithms for optimal motion planning
S. Karaman and E. Frazzoli · 2011
Cited alongside, same era.
Exploration in model-based reinforcement learning by empirically estimating learning progress
M. Lopes, T. Lang, M. Toussaint, and P.Y. Oudeyer · 2012
Cited alongside, same era.
Using trajectory data to improve bayesian optimization for reinforcement learning
A. Wilson, A. Fern, and P. Tadepalli · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, et al · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Stochastic gradient descent as approximate bayesian inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Later among the works it cites.
PRM-RL: Long-range robotic navigation tasks by combining reinforcement learning and sampling-based planning
A. Faust, K. Oslund, O. Ramirez, A. Francis, L. Tapia, M. Fiser, and J. Davidson · 2018
Later among the works it cites.
DORA the explorer: Directed outreaching reinforcement action-selection
L. Fox, L. Choshen, and Y. Loewenstein · 2018
Later among the works it cites.
Bayesian RL for goal-only rewards
P. Morere and F. Ramos · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Learning grounded finite-state representations from unstructured demonstrations
S. Niekum, S. Osentoski, G. Konidaris, S. Chitta, et al · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, and K. Ziebaand · 2016
Cited alongside, same era.
OpenAI Gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Later among the works it cites.
Lectures on Convex Optimization
Y. Nesterov · 2018
Later among the works it cites.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Later among the works it cites.
Parameter space noise for exploration
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz · 2018
Later among the works it cites.
Behavioral cloning from observation
F. Torabi, G. Warnell, and P. Stone · 2018
Later among the works it cites.
Bayesian functional optimization
N. A. Vien, H. Zimmermann, and M. Toussaint · 2018
Later among the works it cites.
Learning to plan via neural exploration-exploitation trees
B. Chen, B. Dai, and L. Song · 2019
Later among the works it cites.
RL-RRT: Kinodynamic motion planning via learning reachability estimators from RL policies
H. Lewis Chiang, J. Hsu, M. Fiser, L. Tapia, and A. Faust · 2019
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J Huizinga, J Lehman, K. O. Stanley, and J. Clune · 2019
Later among the works it cites.
Probabilistic completeness of RRT for geometric and kinodynamic planning with forward propagation
M. Kleinbort, K. Solovey, Z. Littlefield, K. E. Bekris, and D. Halperin · 2019
Later among the works it cites.