Fetching the paper…
Reading the bibliography…
Inverse reinforcement learning (IRL) is the problem of inferring the reward function of an agent, given its policy or observed behavior.
T. S. Reddy, V. Gopikrishna, G. Zaruba, M. Huber, Inverse reinforcement learning for decentralized non-cooperative multiagent systems, in: 2012 IEEE International Conference on Systems, Man, and Cybernetics (SMC), 2012, pp. 1930–1935
1935
Earlier work this paper cites.
R. Eberhart, J. Kennedy, Particle swarm optimization, in: Proceedings of the IEEE international conference on neural networks, Vol. 4, Citeseer, 1995, pp. 1942–1948
1948
Earlier work this paper cites.
E. T. Jaynes, Information theory and statistical mechanics, Phys. Rev. 106 (1957) 620–630
1957
Earlier work this paper cites.
S. Kullback, Information theory and statistics (1968)
1968
Earlier work this paper cites.
R. Fletcher, Practical methods of optimization, Wiley-Interscience publication, Wiley, 1987
1987
Earlier work this paper cites.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, 1st Edition, John Wiley & Sons, Inc., New York, NY, USA, 1994
1994
Earlier work this paper cites.
M. L. Littman, Markov games as a framework for multi-agent reinforcement learning, in: Proceedings of the eleventh international conference on machine learning, Vol. 157, 1994, pp. 157–163
1994
Earlier work this paper cites.
S. P. Boyd, L. El Ghaoui, E. Feron, V. Balakrishnan, Linear matrix inequalities in system and control theory, SIAM 37 (3) (1995) 479–481
1995
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, A. W. Moore, Reinforcement learning: A survey, J. Artif. Int. Res. 4 (1) (1996) 237–285
1996
Earlier work this paper cites.
J. Choi, K. eung Kim, Map inference for bayesian inverse reinforcement learning, in: Advances in Neural Information Processing Systems 24, 2011, pp. 1989–1997
1997
Earlier work this paper cites.
S. Russell, Learning agents for uncertain environments (extended abstract), in: Proceedings of the Eleventh Annual Conference on Computational Learning Theory, COLT’ 98, ACM, New York, NY, USA, 1998, pp. 101–103
1998
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, A. R. Cassandra, Planning and acting in partially observable stochastic domains, Artif. Intell. 101 (1-2) (1998) 99–134
1998
Earlier work this paper cites.
C. Boutilier, Sequential optimality and coordination in multiagent systems, in: Proceedings of the 16th International Joint Conference on Artifical Intelligence - Volume 1, IJCAI’99, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 1999, pp. 478–485
1999
Earlier work this paper cites.
A. Ng, S. Russell, Algorithms for inverse reinforcement learning, Proceedings of the Seventeenth International Conference on Machine Learning 0 (2000) 663–670
2000
Earlier work this paper cites.
L. Peshkin, K.-E. Kim, N. Meuleau, L. P. Kaelbling, Learning to cooperate via policy search, in: Proceedings of the 16th Conference on Uncertainty in Artificial Intelligence, UAI ’00, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2000, pp. 489–496
2000
Earlier work this paper cites.
R. Malouf, A comparison of algorithms for maximum entropy parameter estimation, in: Proceedings of the 6th Conference on Natural Language Learning - Volume 20, COLING-02, Association for Computational Linguistics, Stroudsburg, PA, USA, 2002, pp. 1–7
2002
Earlier work this paper cites.
S. Wang, R. Rosenfeld, Y. Zhao, D. Schuurmans, The Latent Maximum Entropy Principle, in: IEEE International Symposium on Information Theory, 2002, pp. 131–131
2002
Earlier work this paper cites.
D. V. Pynadath, M. Tambe, The communicative multiagent team decision problem: Analyzing teamwork theories and models, J. Artif. Int. Res. 16 (1) (2002) 389–423
2002
Earlier work this paper cites.
D. S. Bernstein, R. Givan, N. Immerman, S. Zilberstein, The complexity of decentralized control of markov decision processes, Math. Oper. Res. 27 (4) (2002) 819–840
2002
Earlier work this paper cites.
S. Russell, P. Norvig, Artificial Intelligence: A Modern Approach (Second Edition), Prentice Hall, 2003
2003
Earlier work this paper cites.
P. Abbeel, A. Y. Ng, Apprenticeship learning via inverse reinforcement learning, in: Proceedings of the Twenty-first International Conference on Machine Learning, ICML ’04, ACM, New York, NY, USA, 2004, pp. 1–8
2004
Earlier work this paper cites.
arXiv:0410076v1
P. D. Grünwald, A. P. Dawid, Game theory, maximum entropy, minimum discrepancy and robust bayesian decision theory, The Annals of Statistics 32 (1) (2004) 1367–1433 · 2004
Earlier work this paper cites.
A. Y. Ng, Feature selection, l1 vs. l2 regularization, and rotational invariance, in: Proceedings of the Twenty-first International Conference on Machine Learning, ICML ’04, ACM, New York, NY, USA, 2004, pp. 78–
2004
Earlier work this paper cites.
B. Taskar, V. Chatalbashev, D. Koller, C. Guestrin, Learning structured prediction models: A large margin approach, in: 22nd International Conference on Machine Learning, 2005, p. 896–903
2005
Earlier work this paper cites.
P. J. Gmytrasiewicz, P. Doshi, A framework for sequential planning in multi-agent settings, J. Artif. Int. Res. 24 (1) (2005) 49–79
2005
Earlier work this paper cites.
P. Abbeel, A. Coates, M. Quigley, A. Y. Ng, An application of reinforcement learning to aerobatic helicopter flight, in: Proceedings of the 19th International Conference on Neural Information Processing Systems, NIPS’06, MIT Press, Cambridge, MA, USA, 2006, pp. 1–8
2006
Earlier work this paper cites.
N. D. Ratliff, J. A. Bagnell, M. A. Zinkevich, Maximum margin planning, in: Proceedings of the 23rd International Conference on Machine Learning, ICML ’06, ACM, New York, NY, USA, 2006, pp. 729–736
2006
Earlier work this paper cites.
N. Ratliff, D. Bradley, J. A. Bagnell, J. Chestnutt, Boosting structured prediction for imitation learning, in: Proceedings of the 19th International Conference on Neural Information Processing Systems, NIPS’06, MIT Press, Cambridge, MA, USA, 2006, pp. 1153–1160
2006
Earlier work this paper cites.
2007
Cited alongside, same era.
U. Syed, R. E. Schapire, A game-theoretic approach to apprenticeship learning, in: Proceedings of the 20th International Conference on Neural Information Processing Systems, NIPS’07, Curran Associates Inc., USA, 2007, pp. 1449–1456
2007
Cited alongside, same era.
D. Ramachandran, E. Amir, Bayesian inverse reinforcement learning, in: Proceedings of the 20th International Joint Conference on Artifical Intelligence, IJCAI’07, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2007, pp. 2586–2591
2007
Cited alongside, same era.
U. Syed, R. E. Schapire, A Game-Theoretic Approach to Apprenticeship Learning—Supplement (2007)
2007
Cited alongside, same era.
N. Aghasadeghi, T. Bretl, Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals, in: 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2011, pp. 1561–1566
2011
Later among the works it cites.
A. Boularias, J. Kober, J. Peters, Relative entropy inverse reinforcement learning, in: Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2011, Fort Lauderdale, USA, April 11-13, 2011, 2011, pp. 182–189
2011
Later among the works it cites.
S. Levine, Z. Popović, V. Koltun, Nonlinear inverse reinforcement learning with gaussian processes, in: Proceedings of the 24th International Conference on Neural Information Processing Systems, NIPS’11, Curran Associates Inc., USA, 2011, pp. 19–27
2011
Later among the works it cites.
M. Babes-Vroman, V. Marivate, K. Subramanian, M. Littman, Apprenticeship learning about multiple intentions, in: 28th International Conference on Machine Learning, ICML 2011, 2011, pp. 897–904
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Z. Kolter, P. Abbeel, A. Y. Ng, Hierarchical apprenticeship learning, with application to quadruped locomotion, in: Proceedings of the 20th International Conference on Neural Information Processing Systems, NIPS’07, Curran Associates Inc., USA, 2007, pp. 769–776
2007
Cited alongside, same era.
B. D. Ziebart, A. Maas, J. A. Bagnell, A. K. Dey, Maximum entropy inverse reinforcement learning, in: Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 3, AAAI’08, AAAI Press, 2008, pp. 1433–1438
2008
Cited alongside, same era.
B. D. Ziebart, A. L. Maas, A. K. Dey, J. A. Bagnell, Navigate like a cabbie: Probabilistic reasoning from observed context-aware behavior, in: Proceedings of the 10th International Conference on Ubiquitous Computing, UbiComp ’08, ACM, New York, NY, USA, 2008, pp. 322–331
2008
Cited alongside, same era.
A. S. David Silver, James Bagnell, High performance outdoor navigation from overhead data using imitation learning, in: Robotics: Science and Systems IV, Zurich, Switzerland, 2008
2008
Cited alongside, same era.
A. Coates, P. Abbeel, A. Y. Ng, Learning for control from multiple demonstrations, in: Proceedings of the 25th International Conference on Machine Learning, ICML ’08, ACM, New York, NY, USA, 2008, pp. 144–151
2008
Cited alongside, same era.
M. T. J. Spaan, F. S. Melo, Interaction-driven markov games for decentralized multiagent planning under uncertainty, in: Proceedings of the 7th International Joint Conference on Autonomous Agents and Multiagent Systems - Volume 1, AAMAS ’08, International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 2008, pp. 525–532
2008
Cited alongside, same era.
A. Coates, P. Abbeel, A. Y. Ng, Apprenticeship learning for helicopter control, Communications of the ACM 52 (7) (2009) 97–105
2009
Cited alongside, same era.
B. D. Argall, S. Chernova, M. Veloso, B. Browning, A survey of robot learning from demonstration, Robot. Auton. Syst. 57 (5) (2009) 469–483
2009
Cited alongside, same era.
2011
Later among the works it cites.
A. Vogel, D. Ramachandran, R. Gupta, A. Raux, Improving hybrid vehicle fuel efficiency using inverse reinforcement learning, in: AAAI Conference on Artificial Intelligence, 2012
2012
Later among the works it cites.
A. Boularias, O. Krömer, J. Peters, Structured apprenticeship learning, in: P. A. Flach, T. De Bie, N. Cristianini (Eds.), Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2012, Bristol, UK, September 24-28, 2012. Proceedings, Part II, Springer Berlin Heidelberg, Berlin, Heidelberg, 2012, pp. 227–242
2012
Later among the works it cites.
E. Klein, M. Geist, B. Piot, O. Pietquin, Inverse reinforcement learning through structured classification, in: Proceedings of the 25th International Conference on Neural Information Processing Systems, NIPS’12, Curran Associates Inc., USA, 2012, pp. 1007–1015
2012
Later among the works it cites.
C. Dimitrakakis, C. A. Rothkopf, Bayesian multitask inverse reinforcement learning, in: Proceedings of the 9th European Conference on Recent Advances in Reinforcement Learning, EWRL’11, Springer-Verlag, Berlin, Heidelberg, 2012, pp. 273–284
2012
Later among the works it cites.
P. Vernaza, J. A. Bagnell, Efficient high-dimensional maximum entropy modeling via symmetric partition functions, in: Proceedings of the 25th International Conference on Neural Information Processing Systems, NIPS’12, Curran Associates Inc., USA, 2012, pp. 575–583
2012
Later among the works it cites.
K. M. Kitani, B. D. Ziebart, J. A. Bagnell, M. Hebert, Activity forecasting, in: Proceedings of the 12th European Conference on Computer Vision - Volume Part IV, ECCV’12, Springer-Verlag, Berlin, Heidelberg, 2012, pp. 201–214
2012
Later among the works it cites.
J. Choi, K.-E. Kim, Nonparametric bayesian inverse reinforcement learning for multiple reward functions, in: Proceedings of the 25th International Conference on Neural Information Processing Systems, NIPS’12, Curran Associates Inc., USA, 2012, pp. 305–313
2012
Later among the works it cites.
E. Klein, B. Piot, M. Geist, O. Pietquin, A cascaded supervised learning approach to inverse reinforcement learning, in: Proceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases - Volume 8188, ECML PKDD 2013, Springer-Verlag New York, Inc., New York, NY, USA, 2013, pp. 1–16
2013
Later among the works it cites.
C. A. Rothkopf, D. H. Ballard, Modular inverse reinforcement learning for visuomotor behavior, Biol. Cybern. 107 (4) (2013) 477–490
2013
Later among the works it cites.
J. Choi, K.-E. Kim, Bayesian nonparametric feature construction for inverse reinforcement learning, in: Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, IJCAI ’13, AAAI Press, 2013, pp. 1287–1293
2013
Later among the works it cites.
M. Kalakrishnan, P. Pastor, L. Righetti, S. Schaal, Learning objective functions for manipulation, in: IEEE International Conference on Robotics and Automation (ICRA), 2013, 2013, pp. 1331–1336
2013
Later among the works it cites.
M. C. Vroman, MAXIMUM LIKELIHOOD INVERSE REINFORCEMENT LEARNING, Ph.D. thesis, Rutgers, The State University of New Jersey (2014)
2014
Later among the works it cites.
S. Levine, P. Abbeel, Learning neural network policies with guided policy search under unknown dynamics, in: Proceedings of the 27th International Conference on Neural Information Processing Systems, NIPS’14, MIT Press, Cambridge, MA, USA, 2014, pp. 1071–1079
2014
Later among the works it cites.
M. Kuderer, S. Gulati, W. Burgard, Learning driving styles for autonomous vehicles from demonstration, in: IEEE International Conference on Robotics and Automation (ICRA), 2015, pp. 2641–2646
2015
Later among the works it cites.
K. Bogert, P. Doshi, Multi-robot inverse reinforcement learning under occlusion with state transition estimation, in: Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems, AAMAS ’15, International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 2015, pp. 1837–1838
2015
Later among the works it cites.
T. Munzer, B. Piot, M. Geist, O. Pietquin, M. Lopes, Inverse reinforcement learning in relational domains, in: Proceedings of the 24th International Conference on Artificial Intelligence, IJCAI’15, AAAI Press, 2015, pp. 3735–3741
2015
Later among the works it cites.
K. Bogert, P. Doshi, Toward estimating others’ transition models under occlusion for multi-robot irl, in: Proceedings of the 24th International Conference on Artificial Intelligence, IJCAI’15, AAAI Press, 2015, pp. 1867–1873
2015
Later among the works it cites.
doi:10.1177/0278364915619772
H. Kretzschmar, M. Spies, C. Sprunk, W. Burgard, Socially compliant mobile robot navigation via inverse reinforcement learning, The International Journal of Robotics Research 35 (11) (2016) 1289–1307 · 2016
Later among the works it cites.
B. Kim, J. Pineau, Socially adaptive path planning in human environments using inverse reinforcement learning, International Journal of Social Robotics 8 (1) (2016) 51–66
2016
Later among the works it cites.
K. Shiarlis, J. Messias, S. Whiteson, Inverse reinforcement learning from failure, in: Proceedings of the 2016 International Conference on Autonomous Agents and Multiagent Systems, AAMAS ’16, International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 2016, pp. 1060–1068
2016
Later among the works it cites.
K. Bogert, J. F.-S. Lin, P. Doshi, D. Kulic, Expectation-maximization for inverse reinforcement learning with hidden data, in: Proceedings of the 2016 International Conference on Autonomous Agents and Multiagent Systems, AAMAS ’16, International Foundation for Autonomous Agents and Multiagent Systems, 2016, pp. 1034–1042
2016
Later among the works it cites.
D. S. Brown, S. Niekum, Efficient probabilistic performance bounds for inverse reinforcement learning, in: Thirty-Second AAAI Conference on Artificial Intelligence, 2018
2018
Closest in time.
D. Brown, W. Goo, P. Nagarajan, S. Niekum, Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations, in: Proceedings of the 36th International Conference on Machine Learning, Vol. 97 of Proceedings of Machine Learning Research, 2019, pp. 783–792
2019
Closest in time.