Fetching the paper…
Reading the bibliography…
Learning robotic manipulation tasks using reinforcement learning with sparse rewards is currently impractical due to the outrageous data requirements.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning , vol. 8, no. 3-4, pp. 279–292, 1992
1992
Earlier work this paper cites.
J. L. Elman, “Learning and development in neural networks: the importance of starting small,” Cognition , vol. 48, no. 1, p. 71–99, 1993
1993
Earlier work this paper cites.
P. Dayan and G. E. Hinton, “Feudal reinforcement learning,” in Advances in neural information processing systems , 1993, pp. 271–278
1993
Earlier work this paper cites.
T. G. Dietterich, “The maxq method for hierarchical reinforcement learning,” in In Proceedings of the Fifteenth International Conference on Machine Learning . Morgan Kaufmann, 1998, pp. 118–126
1998
Earlier work this paper cites.
R. Sutton, D. Precup, and S. Singh, “Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning,” Artificial Intelligence , vol. 112, pp. 181–211, 1999
1999
Earlier work this paper cites.
L. P. Kaelbling, T. Oates, N. Hernandez, and S. Finney, “Learning in worlds with objects,” in Working Notes of the AAAI Stanford Spring Symposium on Learning Grounded Representations , 2001, pp. 31–36
2001
Earlier work this paper cites.
S. Džeroski, L. De Raedt, and K. Driessens, “Relational reinforcement learning,” Machine Learning , vol. 43, no. 1, p. 7–52, Apr 2001
2001
Earlier work this paper cites.
M. Van Otterlo, “Relational representations in reinforcement learning: Review and open problems,” in Proceedings of the ICML , vol. 2, 2002
2002
Earlier work this paper cites.
P.-Y. Oudeyer and F. Kaplan, “What is intrinsic motivation? a typology of computational approaches,” Frontiers in neurorobotics , 2009
2009
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the 26th annual international conference on machine learning . ACM, 2009, pp. 41–48
2009
Earlier work this paper cites.
S. Russell and P. Norvig, Artificial Intelligence: A Modern Approach , 3rd ed. Prentice Hall Press, 2009
2009
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the 26th Annual International Conference on Machine Learning , ser. ICML ’09. New York, NY, USA: ACM, 2009, pp. 41–48. [Online]. Available: http://doi.acm.org/10.1145/1553374.1553380
2009
Earlier work this paper cites.
P. Kormushev, S. Calinon, and D. G. Caldwell, “Robot motor skill coordination with em-based reinforcement learning,” in 2010 IEEE/RSJ international conference on intelligent robots and systems . IEEE, 2010, pp. 3232–3237
2010
Earlier work this paper cites.
J. Schmidhuber, “Formal theory of creativity, fun, and intrinsic motivation (1990–2010),” IEEE Transactions on Autonomous Mental Development , 2010
2010
Earlier work this paper cites.
M. P. Deisenroth, C. E. Rasmussen, and D. Fox, “Learning to control a low-cost manipulator using data-efficient reinforcement learning,” Robotics Science and Systems , pp. 57–64, 2011
2011
Earlier work this paper cites.
M. Lopes, T. Lang, M. Toussaint, and P.-Y. Oudeyer, “Exploration in model-based reinforcement learning by empirically estimating learning progress,” in NIPS , 2012
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control.” in IROS . IEEE, 2012, pp. 5026–5033. [Online]. Available: http://dblp.uni-trier.de/db/conf/iros/iros2012.html#TodorovET12
2012
Earlier work this paper cites.
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell, “Decaf: A deep convolutional activation feature for generic visual recognition,” in International conference on machine learning , 2014, pp. 647–655
2014
Earlier work this paper cites.
J. Bruna, W. Zaremba, A. Szlam, and Y. Lecun, “Spectral networks and locally connected networks on graphs,” in International Conference on Learning Representations (ICLR2014), CBLS, April 2014 , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems , 2015, pp. 91–99
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Toussaint, “Logic-geometric programming: An optimization-based approach to combined task and motion planning,” in Twenty-Fourth International Joint Conference on Artificial Intelligence , 2015
2015
Cited alongside, same era.
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” in International Conference on Machine Learning , 2015, pp. 1312–1320
2015
Cited alongside, same era.
2016
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” JMLR , 2016
2016
Cited alongside, same era.
P. Agrawal, A. Nair, P. Abbeel, J. Malik, and S. Levine, “Learning to poke by poking: Experiential learning of intuitive physics,” NIPS , 2016
2017
Later among the works it cites.
Y. Duan, M. Andrychowicz, B. Stadie, O. J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba, “One-shot imitation learning,” in Advances in neural information processing systems , 2017, pp. 1087–1098
2017
Later among the works it cites.
2017
Later among the works it cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba, “Hindsight experience replay,” in Advances in Neural Information Processing Systems 30 , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds. Curran Associates, Inc., 2017, pp. 5048–5058. [Online]. Available: http://papers.nips.cc/paper/7090-hindsight-experience-replay.pdf
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
M. Defferrard, X. Bresson, and P. Vandergheynst, “Convolutional neural networks on graphs with fast localized spectral filtering,” in Advances in Neural Information Processing Systems 29 , D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett, Eds. Curran Associates, Inc., 2016, pp. 3844–3852. [Online]. Available: http://papers.nips.cc/paper/6081-convolutional-neural-networks-on-graphs-with-fast-localized-spectral-filtering.pdf
2016
Cited alongside, same era.
2016
Cited alongside, same era.
P. Battaglia, R. Pascanu, M. Lai, D. J. Rezende, and K. kavukcuoglu, “Interaction networks for learning about objects, relations and physics,” in Proceedings of the 30th International Conference on Neural Information Processing Systems , ser. NIPS’16. USA: Curran Associates Inc., 2016, pp. 4509–4517. [Online]. Available: http://dl.acm.org/citation.cfm?id=3157382.3157601
2016
Cited alongside, same era.
2016
Cited alongside, same era.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” ICLR , 2016
2016
Cited alongside, same era.
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. Dragan, “Inverse reward design,” in Advances in neural information processing systems , 2017, pp. 6765–6774
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Later among the works it cites.
2018
Later among the works it cites.
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver, “Rainbow: Combining improvements in deep reinforcement learning,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus, “Intrinsic motivation and automatic curricula via asymmetric self-play,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=SkT5Yg-RZ
2018
Later among the works it cites.
T. Wang, R. Liao, J. Ba, and S. Fidler, “Nervenet: Learning structured policy with graph neural networks,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=S1sqHMZCb
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
I. Popov, N. Heess, T. P. Lillicrap, R. Hafner, G. Barth-Maron, M. Vecerik, T. Lampe, T. Erez, Y. Tassa, and M. Riedmiller, “Data-efficient deep reinforcement learning for dexterous manipulation,” 2018. [Online]. Available: https://openreview.net/forum?id=SJdCUMZAW
2018
Later among the works it cites.
O. Kroemer, S. Leischnig, S. Luettgen, and J. Peters, “A kernel-based approach to learning contact distributions for robot manipulation tasks,” Autonomous Robots , vol. 42, no. 3, pp. 581–600, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction , 2nd ed. The MIT Press, 2018. [Online]. Available: http://incompleteideas.net/book/the-book-2nd.html
2018
Later among the works it cites.
V. Zambaldi, D. Raposo, A. Santoro, V. Bapst, Y. Li, I. Babuschkin, K. Tuyls, D. Reichert, T. Lillicrap, E. Lockhart, et al. , “Deep reinforcement learning with relational inductive biases,” International Conference on Learning Representations , 2019
2019
Closest in time.
2019
Closest in time.
M. Janner, S. Levine, W. T. Freeman, J. B. Tenenbaum, C. Finn, and J. Wu, “Reasoning about physical interactions with object-oriented prediction and planning,” in International Conference on Learning Representations , 2019
2019
Closest in time.