Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has achieved tremendous success as a general framework for learning how to make decisions.
Ranking-Based Reward Extrapolation without Rankings
Daniel S. Brown, Wonjoon Goo, and Scott Niekum. 2019 · 1907
Earlier work this paper cites.
Theories of bounded rationality
Herbert A Simon. 1972 · 1972
Earlier work this paper cites.
Decision making in action: Models and methods.. In This book is an outcome of a workshop held in Dayton, OH, Sep 25–27, 1989. Ablex Publishing
Gary A Klein, Judith Ed Orasanu, Roberta Ed Calderwood, and Caroline E Zsambok. 1993 · 1989
Earlier work this paper cites.
The recognition-primed decision (RPD) model: Looking back, looking forward
Gary Klein. 1997 · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning.. In Icml , Vol. 1. 2
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning . ACM, 1
Pieter Abbeel and Andrew Y Ng. 2004 · 2004
Earlier work this paper cites.
Feature Selection for Medical Data Mining: Comparisons of Expert Judgment and Automatic Approaches. In Proceedings of the 19th IEEE Symposium on Computer-Based Medical Systems (CBMS ’06) . IEEE Computer Society, Washington, DC, USA, 165–170
Tsang-Hsiang Cheng, Chih-Ping Wei, and Vincent S. Tseng. 2006 · 2006
Earlier work this paper cites.
Active learning with feedback on features and instances
Hema Raghavan, Omid Madani, and Rosie Jones. 2006 · 2006
Earlier work this paper cites.
Maximum margin planning. In Proceedings of the 23rd international conference on Machine learning . ACM, 729–736
Nathan D Ratliff, J Andrew Bagnell, and Martin A Zinkevich. 2006 · 2006
Earlier work this paper cites.
Bayesian Inverse Reinforcement Learning.. In IJCAI , Vol. 7. 2586–2591
Deepak Ramachandran and Eyal Amir. 2007 · 2007
Earlier work this paper cites.
Collaborative filtering recommender systems
J Ben Schafer, Dan Frankowski, Jon Herlocker, and Shilad Sen. 2007 · 2007
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 3 (Chicago, Illinois) (AAAI’08) . AAAI Press, 1433–1438
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey. 2008 · 2008
Earlier work this paper cites.
Learning table tennis with a mixture of motor primitives. In Proceedings of the International Conference on Humanoid Robots (ICHR) . IEEE, 411–416
Katharina Müelling, Jens Kober, and Jan Peters. 2010 · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D. Ziebart. 2010 · 2010
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 15) , Geoffrey Gordon, David Dunson, and Miroslav Dudík (Eds.). PMLR, Fort Lauderdale, FL, USA, 627–635
Stephane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
Nonparametric Bayesian Inverse Reinforcement Learning for Multiple Reward Functions
Jaedeug Choi and Kee eung Kim. 2012 · 2012
Earlier work this paper cites.
Bayesian Multitask Inverse Reinforcement Learning. In Proceedings of the 9th European Conference on Recent Advances in Reinforcement Learning (Athens, Greece) (EWRL’11) . Springer-Verlag, Berlin, Heidelberg, 273–284
Christos Dimitrakakis and Constantin A. Rothkopf. 2012 · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 5026–5033
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Cited alongside, same era.
Learning to select and generalize striking movements in robot table tennis
Katharina Mülling, Jens Kober, Oliver Kroemer, and Jan Peters. 2013 · 2013
Cited alongside, same era.
Towards robot skill learning: From simple skills to table tennis. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD) . Springer, 627–631
Jan Peters, Jens Kober, Katharina Mülling, Oliver Krämer, and Gerhard Neumann. 2013 · 2013
Cited alongside, same era.
The algorithmic anatomy of model-based evaluation
Nathaniel D Daw and Peter Dayan. 2014 · 2014
Cited alongside, same era.
Learning strategies in table tennis using inverse reinforcement learning
Katharina Muelling, Abdeslam Boularias, Betty Mohler, Bernhard Schölkopf, and Jan Peters. 2014 · 2014
InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
Yunzhu Li, Jiaming Song, and Stefano Ermon. 2017 · 2017
Later among the works it cites.
Game-Theoretic Modeling of Human Adaptation in Human-Robot Collaboration. In Proceedings of the 2017 ACM/IEEE International Conference on Human-Robot Interaction (Vienna, Austria) (HRI ’17) . ACM, New York, NY, USA, 323–331
Stefanos Nikolaidis, Swaprava Nath, Ariel D. Procaccia, and Siddhartha Srinivasa. 2017 · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Yee Teh, Victor Bapst, Wojciech M. Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu. 2017 · 2017
Later among the works it cites.
Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In International Conference on Learning Representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient Model Learning for Human-Robot Collaborative Tasks
Stefanos Nikolaidis, Keren Gu, Ramya Ramakrishnan, and Julie A. Shah. 2014 · 2014
Cited alongside, same era.
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Cited alongside, same era.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell. 2015 · 2015
Cited alongside, same era.
Maximum entropy deep inverse reinforcement learning
Markus Wulfmeier, Peter Ondruska, and Ingmar Posner. 2015 · 2015
Cited alongside, same era.
Concrete Problems in AI Safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization. In International Conference on Machine Learning . 49–58
Chelsea Finn, Sergey Levine, and Pieter Abbeel. 2016 · 2016
Cited alongside, same era.
Justin Fu, Katie Luo, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Human-machine collaborative optimization via apprenticeship scheduling
Matthew Gombolay, Reed Jensen, Jessica Stigile, Toni Golen, Neel Shah, Sung-Hyun Son, and Julie Shah. 2018a · 2018
Later among the works it cites.
Robotic assistance in the coordination of patient care
Matthew Gombolay, Xi Jessie Yang, Bradley Hayes, Nicole Seo, Zixi Liu, Samir Wadhwania, Tania Yu, Neel Shah, Toni Golen, and Julie Shah. 2018b · 2018
Later among the works it cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80) , Jennifer Dy and Andreas Krause (Eds.). PMLR, Stockholmsmässan, Stockholm Sweden, 1861–1870
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning. In Thirty-Second AAAI Conference on Artificial Intelligence
Peter Henderson, Wei-Di Chang, Pierre-Luc Bacon, David Meger, Joelle Pineau, and Doina Precup. 2018 · 2018
Later among the works it cites.
Online optimal trajectory generation for robot table tennis
Okan Koç, Guilherme Maeda, and Jan Peters. 2018 · 2018
Later among the works it cites.
Preference Learning in Assistive Robotics: Observational Repeated Inverse Reinforcement Learning. In Proceedings of the 3rd Machine Learning for Healthcare Conference (Proceedings of Machine Learning Research, Vol. 85) , Finale Doshi-Velez, Jim Fackler, Ken Jung, David Kale, Rajesh Ranganath, Byron Wallace, and Jenna Wiens (Eds.). PMLR, Palo Alto, California, 420–439
Bryce Woodworth, Francesco Ferrari, Teofilo E. Zosa, and Laurel D. Riek. 2018 · 2018
Later among the works it cites.
Distilling Policy Distillation. In Proceedings of Machine Learning Research (Proceedings of Machine Learning Research, Vol. 89) , Kamalika Chaudhuri and Masashi Sugiyama (Eds.). PMLR, 1331–1340
Wojciech M. Czarnecki, Razvan Pascanu, Simon Osindero, Siddhant Jayakumar, Grzegorz Swirszcz, and Max Jaderberg. 2019 · 2019
Later among the works it cites.
Diversity is All You Need: Learning Skills without a Reward Function. In International Conference on Learning Representations
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. 2019 · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Learning a Prior over Intent via Meta-Inverse Reinforcement Learning. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, Long Beach, California, USA, 6952–6962
Kelvin Xu, Ellis Ratner, Anca Dragan, Sergey Levine, and Chelsea Finn. 2019 · 2019
Later among the works it cites.
Learning Novel Policies For Tasks. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, Long Beach, California, USA, 7483–7492
Yunbo Zhang, Wenhao Yu, and Greg Turk. 2019 · 2019
Later among the works it cites.