Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) can automate a wide variety of robotic skills, but learning each new skill requires considerable real-world data collection and manual representation engineering to design policy classes or features.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3-4, pp. 229–256, May 1992. [Online]. Available: http://dx.doi.org/10.1007/BF00992696
1992
Earlier work this paper cites.
R. Caruana, “Learning many related tasks at the same time with backpropagation,” in In Advances in Neural Information Processing Systems 7 . Morgan Kaufmann, 1995, pp. 657–664
1995
Earlier work this paper cites.
——, “Multitask learning,” Machine Learning , vol. 28, no. 1, pp. 41–75, 1997
1997
Earlier work this paper cites.
C. Drummond, “Accelerating reinforcement learning by composing solutions of automatically identified subtasks,” J. Artif. Intell. Res. (JAIR) , 2002
2002
Earlier work this paper cites.
C. Drummond, “Accelerating reinforcement learning by composing solutions of automatically identified subtasks,” JAIR , vol. 16, pp. 59–104, 2002. [Online]. Available: http://jair.org/papers/paper904.html
2002
Earlier work this paper cites.
C. Guestrin, D. Koller, C. Gearhart, and N. Kanodia, “Generalizing plans to new environments in relational mdps,” in In International Joint Conference on Artificial Intelligence , 2003
2003
Earlier work this paper cites.
M. G. Madden and T. Howley, “Transfer of experience between reinforcement learning environments with progressive difficulty,” Artificial Intelligence Review , vol. 21, no. 3, 2004
2004
Earlier work this paper cites.
G. Konidaris and A. Barto, “Autonomous shaping: knowledge transfer in reinforcement learning,” in International Conference on Machine Learning (ICML) , 2006, pp. 489–496
2006
Earlier work this paper cites.
J. Ramon, K. Driessens, and T. Croonenborghs, Transfer Learning in Reinforcement Learning Problems Through Partial Policy Recycling , 2007
2007
Earlier work this paper cites.
G. Konidaris and A. G. Barto, “Building portable options: Skill transfer in reinforcement learning.” in Proc. International Joint Conference on Artificial Intelligence , 2007, pp. 895–900
2007
Earlier work this paper cites.
M. Taylor, P. Stone, and Y. Liu, “Transfer learning via inter-task mappings for temporal difference learning,” Journal of Machine Learning Research , vol. 8, no. 1, pp. 2125–2167, 2007
2007
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural Networks , 2008
2008
Cited alongside, same era.
M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research , vol. 10, pp. 1633–1685, 2009
2009
Cited alongside, same era.
L. Mihalkova and R. J. Mooney, “Transfer learning from minimal target data by mapping across relational domains,” in Transfer Learning from Minimal Target Data by Mapping across Relational Domains , 2009
2009
Cited alongside, same era.
J. Peters, K. Mülling, and Y. Altun, “Relative entropy policy search,” in Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2010, Atlanta, Georgia, USA, July 11-15, 2010 , 2010
2010
Cited alongside, same era.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in Advances in Neural Information Processing Systems , 2014
2014
Later among the works it cites.
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, “Trust region policy optimization,” in International Conference on Machine Learning (ICML) , 2015
2015
Later among the works it cites.
E. Tzeng, J. Hoffman, T. Darrell, and K. Saenko, “Simultaneous deep transfer across domains and tasks,” in International Conference in Computer Vision (ICCV) , 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2012
Cited alongside, same era.
2013
Cited alongside, same era.
M. P. Deisenroth, G. Neumann, and J. Peters, “A survey on policy search for robotics,” Found. Trends Robot , vol. 2, no. 1–2, pp. 1–142, Aug. 2013
2013
Cited alongside, same era.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” I. J. Robotics Res. , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Cited alongside, same era.
2013
Cited alongside, same era.
H. B. Ammar, E. Eaton, P. Ruvolo, and M. E. Taylor, “Online multi-task learning for policy gradient methods,” Journal of Machine Learning Research , 2014
2014
Cited alongside, same era.
A. K. I. S. R. S. Nitish Srivastava, Geoffrey Hinton, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research , vol. 15, 2014
2014
Cited alongside, same era.
J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing , vol. 71
Cited in the paper.
2015
Later among the works it cites.
2015
Later among the works it cites.
I. Mordatch, N. Mishra, C. Eppner, and P. Abbeel, “Combining model-based policy search with online model learning for control of physical humanoids,” in 2016 IEEE International Conference on Robotics and Automation, ICRA 2016, Stockholm, Sweden, May 16-21, 2016 , 2016, pp. 242–248
2016
Closest in time.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies.” Journal of Machine Learning Research , vol. 17, pp. 1–40, 2016
2016
Closest in time.
2016
Closest in time.