Fetching the paper…
Reading the bibliography…
Exploration in environments with sparse rewards has been a persistent problem in reinforcement learning (RL).
T. Winograd,
1972
Earlier work this paper cites.
D. A. Pomerleau, “Alvinn: An autonomous land vehicle in a neural network,”
1989
Earlier work this paper cites.
L. Kavraki
1996
Earlier work this paper cites.
S. Schaal, “Robot learning from demonstration,”
1997
Earlier work this paper cites.
A. Ng and S. Russell, “Algorithms for Inverse Reinforcement Learning,”
2000
Earlier work this paper cites.
J. Nakanishi
2004
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in
2004
Earlier work this paper cites.
B. D. Ziebart
2008
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,”
2008
Earlier work this paper cites.
J. Kober and J. Peter, “Policy search for motor primitives in robotics,” in
2008
Earlier work this paper cites.
M. Kalakrishnan
2009
Earlier work this paper cites.
J. Peters, K. Mülling, and Y. Altün, “Relative Entropy Policy Search,”
2010
Earlier work this paper cites.
M. P. Deisenroth, C. E. Rasmussen, and D. Fox, “Learning to Control a Low-Cost Manipulator using Data-Efficient Reinforcement Learning,”
2011
Cited alongside, same era.
S. Ross, G. J. Gordon, and J. A. Bagnell, “A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,” in
2011
Cited alongside, same era.
M. P. Deisenroth and C. E. Rasmussen, “Pilco: A model-based and data-efficient approach to policy search,” in
2011
Cited alongside, same era.
L. P. Kaelbling and T. Lozano-Perez, “Hierarchical task and motion planning in the now,”
2011
Cited alongside, same era.
E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in
2012
Cited alongside, same era.
S. Srivastava
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2015
Later among the works it cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in
2015
Later among the works it cites.
2016
Later among the works it cites.
C. Finn, S. Levine, and P. Abbeel, “Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization,” in
2016
Later among the works it cites.
D. Silver
2016
Later among the works it cites.
2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
A. Giusti
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
J. Schulman
2015
Cited alongside, same era.
T. Schaul
2015
Cited alongside, same era.
Later among the works it cites.
M. Andrychowicz
2017
Closest in time.
2017
Closest in time.
C. Florensa
2017
Closest in time.