Fetching the paper…
Reading the bibliography…
This work handles the inverse reinforcement learning (IRL) problem where only a small number of demonstrations are available from a demonstrator for each high-dimensional task, insufficient to estimate an accurate reward function.
R. Kalman and M. M. C. B. D. R. I. for Advanced Studies. Center for Control Theory, When is a Linear Control System Optimal?. , ser. RIAS technical report. Martin Marietta Corporation, Research Institute for Advanced Studies, Center for Control Theory, 1963
1963
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press Cambridge, 1998, vol. 1, no. 1
1998
Earlier work this paper cites.
A. Y. Ng and S. Russell, “Algorithms for inverse reinforcement learning,” in in Proc. 17th International Conf. on Machine Learning , 2000
2000
Earlier work this paper cites.
R. Vilalta and Y. Drissi, “A perspective view and survey of meta-learning,” Artificial Intelligence Review , vol. 18, no. 2, pp. 77–95, 2002
2002
Earlier work this paper cites.
B. Najafi, K. Aminian, A. Paraschiv-Ionescu, F. Loew, C. J. Bula, and P. Robert, “Ambulatory system for human motion analysis using a kinematic sensor: monitoring of daily physical activity in the elderly,” IEEE Transactions on biomedical Engineering , vol. 50, no. 6, pp. 711–723, 2003
2003
Earlier work this paper cites.
N. Schweighofer and K. Doya, “Meta-learning in reinforcement learning,” Neural Networks , vol. 16, no. 1, pp. 5–9, 2003
2003
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning . ACM, 2004, p. 1
2004
Earlier work this paper cites.
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich, “Maximum margin planning,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006, pp. 729–736
2006
Earlier work this paper cites.
D. Ramachandran and E. Amir, “Bayesian inverse reinforcement learning,” in Proceedings of the 20th International Joint Conference on Artifical Intelligence , ser. IJCAI’07. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2007, pp. 2586–2591
2007
Earlier work this paper cites.
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning,” in Proc. AAAI , 2008, pp. 1433–1438
2008
Earlier work this paper cites.
K. Mombaur, A. Truong, and J.-P. Laumond, “From human to humanoid locomotion—an inverse optimal control approach,” Autonomous robots , vol. 28, no. 3, pp. 369–383, 2010
2010
Earlier work this paper cites.
S. Levine, Z. Popovic, and V. Koltun, “Nonlinear inverse reinforcement learning with gaussian processes,” in Advances in Neural Information Processing Systems 24 , J. Shawe-Taylor, R. S. Zemel, P. L. Bartlett, F. Pereira, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2011, pp. 19–27
2011
Cited alongside, same era.
J. Choi and K.-E. Kim, “Inverse reinforcement learning in partially observable environments,” Journal of Machine Learning Research , vol. 12, no. Mar, pp. 691–730, 2011
2011
Cited alongside, same era.
C. Dimitrakakis and C. A. Rothkopf, “Bayesian multitask inverse reinforcement learning,” in European Workshop on Reinforcement Learning . Springer, 2011, pp. 273–284
2011
Cited alongside, same era.
A. Boularias, J. Kober, and J. R. Peters, “Relative entropy inverse reinforcement learning,” in International Conference on Artificial Intelligence and Statistics , 2011, pp. 182–189
2011
Cited alongside, same era.
2016
Later among the works it cites.
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap, “Meta-learning with memory-augmented neural networks,” in International conference on machine learning , 2016, pp. 1842–1850
2016
Later among the works it cites.
2016
Later among the works it cites.
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, and N. de Freitas, “Learning to learn by gradient descent by gradient descent,” in Advances in Neural Information Processing Systems , 2016, pp. 3981–3989
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
2012
Cited alongside, same era.
2012
Cited alongside, same era.
J. Choi and K.-E. Kim, “Nonparametric bayesian inverse reinforcement learning for multiple reward functions,” in Advances in Neural Information Processing Systems , 2012, pp. 305–313
2012
Cited alongside, same era.
Q. P. Nguyen, B. K. H. Low, and P. Jaillet, “Inverse reinforcement learning with locally consistent reward functions,” in Advances in Neural Information Processing Systems , 2015, pp. 1747–1755
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Later among the works it cites.
2016
Later among the works it cites.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
K. Li and J. W. Burdick, “Bellman Gradient Iteration for Inverse Reinforcement Learning,” ArXiv e-prints , Jul. 2017
2017
Closest in time.