Fetching the paper…
Reading the bibliography…
In this paper we study a new reinforcement learning setting where the environment is non-rewarding, contains several possibly related objects of various controllability, and where an apt agent Bob acts independently, with non-observable intentions.
S. Schaal, “Learning from demonstration,” in Advances in Neural Information Processing Systems 9 . Cambridge, MA: MIT Press, 1997, pp. 1040–1046
1997
Earlier work this paper cites.
A. Coates, P. Abbeel, and A. Y. Ng, “Learning for control from multiple demonstrations,” in Proceedings of the 25th international conference on Machine learning . ACM, 2008, pp. 144–151
2008
Earlier work this paper cites.
B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and autonomous systems , vol. 57, no. 5, pp. 469–483, 2009
2009
Earlier work this paper cites.
A. Stout and A. G. Barto, “Competence progress intrinsic motivation,” in 2010 IEEE 9th International Conference on Development and Learning . IEEE, 2010, pp. 257–262
2010
Earlier work this paper cites.
A. Baranes and P.-Y. Oudeyer, “Intrinsically motivated goal exploration for active motor learning in robots: A case study,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2010) . Taipei, Taiwan, Province Of China: IEEE, 2010
2010
Earlier work this paper cites.
M. Babes, V. Marivate, K. Subramanian, and M. L. Littman, “Apprenticeship learning about multiple intentions,” in Proceedings of the 28th International Conference on Machine Learning (ICML-11) , 2011, pp. 897–904
2011
Earlier work this paper cites.
2011
Earlier work this paper cites.
J. Choi and K.-E. Kim, “Nonparametric bayesian inverse reinforcement learning for multiple reward functions,” in Advances in Neural Information Processing Systems , 2012, pp. 305–313
2012
Earlier work this paper cites.
A. J. Ijspeert, J. Nakanishi, H. Hoffmann, P. Pastor, and S. Schaal, “Dynamical movement primitives: learning attractor models for motor behaviors,” Neural computation , vol. 25, no. 2, pp. 328–373, 2013
2013
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in Proceedings of the 30th International Conference on Machine Learning , 2013, pp. 1–9
2013
Earlier work this paper cites.
A. Baranes and P.-Y. Oudeyer, “Active learning of inverse models with intrinsically motivated goal exploration in robots,” Robotics and Autonomous Systems , vol. 61, no. 1, pp. 49–73, 2013
2013
Earlier work this paper cites.
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” in International Conference on Machine Learning , 2015, pp. 1312–1320
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” in Advances in Neural Information Processing Systems , 2016, pp. 1471–1479
2016
Cited alongside, same era.
2017
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Hausman, Y. Chebotar, S. Schaal, G. Sukhatme, and J. J. Lim, “Multi-modal imitation learning from unstructured demonstrations using generative adversarial nets,” in Advances in Neural Information Processing Systems , 2017, pp. 1235–1245
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Later among the works it cites.
S. Blaes, M. Vlastelica, J.-J. Zhu, and G. Martius, “”control what you can: Intrinsically motivated hierarchical reinforcement learner”,” in Deep RL Workshop NeurIPS, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
B. Kang, Z. Jie, and J. Feng, “Policy optimization with demonstrations,” in International Conference on Machine Learning , 2018, pp. 2474–2483
2018
Later among the works it cites.
N. Duminy, S. M. Nguyen, and D. Duhaut, “Learning a set of interrelated tasks by using sequences of motor policies for a socially guided intrinsically motivated learner,” Frontiers in Neurorobotics , 2019
2019
Closest in time.