Fetching the paper…
Reading the bibliography…
In this paper we consider the problem of robot navigation in simple maze-like environments where the robot has to rely on its onboard sensors to perform the navigation task.
J.-C. Latombe,
1991
Earlier work this paper cites.
P. Dayan, “Improving generalization for temporal difference learning: The successor representation,”
1993
Earlier work this paper cites.
M. Ring, “Continual learning in reinforcement environments,”
1995
Earlier work this paper cites.
R. S. Sutton and A. G. Barto,
1998
Earlier work this paper cites.
S. Thrun, W. Burgard, and D. Fox,
2005
Earlier work this paper cites.
S. M. LaValle,
2006
Earlier work this paper cites.
A. Wilson, A. Fern, S. Ray, and P. Tadepalli, “Multi-task reinforcement learning: a hierarchical Bayesian approach,” in
2007
Earlier work this paper cites.
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup, “Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction.” in
2011
Earlier work this paper cites.
M. E. Taylor and P. Stone, “An introduction to inter-task transfer for reinforcement learning,”
2011
Earlier work this paper cites.
M. Riedmiller, S. Lange, and A. Voigtlaender, “Autonomous reinforcement learning on raw visual input data in a real world application,” in
2012
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,”
2013
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in
2014
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski,
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in
2015
Cited alongside, same era.
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” in
2015
Cited alongside, same era.
R. Jonschkowski and O. Brock, “Learning state representations with robotic priors,”
2015
Cited alongside, same era.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”
H. van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in
2016
Closest in time.
Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas, “Dueling network architectures for deep reinforcement learning,” in
2016
Closest in time.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” in
2016
Closest in time.
E. Parisotto, L. J. Ba, and R. Salakhutdinov, “Actor-mimic: Deep multitask and transfer reinforcement learning,” in
2016
Closest in time.
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell, “Policy distillation,” in
2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in
2016
Cited alongside, same era.
T. D. Kulkarni, A. Saeedi, S. Gautam, and S. J. Gershman, “Deep successor reinforcement learning,”
2016
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel, “Deep spatial autoencoders for visuomotor learning,” in
2016
Closest in time.