Fetching the paper…
Reading the bibliography…
When a person is not satisfied with how a robot performs a task, they can intervene to correct it.
Learning Human Objectives by Evaluating Hypothetical Behavior
Siddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg, and Jan Leike. 2019 · 1912
Earlier work this paper cites.
Theory of games and economic behavior
John Von Neumann and Oskar Morgenstern. 1945 · 1945
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Individual choice behavior
R. Duncan Luce. 1959 · 1959
Earlier work this paper cites.
Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research
Sandra G. Hart and Lowell E. Staveland. 1988 · 1988
Earlier work this paper cites.
Maximum Margin Planning. In Proceedings of the 23rd International Conference on Machine Learning (Pittsburgh, Pennsylvania, USA) (ICML ’06) . Association for Computing Machinery, New York, NY, USA, 729–736
Nathan D. Ratliff, J. Andrew Bagnell, and Martin A. Zinkevich. 2006 · 2006
Earlier work this paper cites.
Goal inference as inverse planning
Chris Baker, Joshua B Tenenbaum, and Rebecca R Saxe. 2007 · 2007
Earlier work this paper cites.
Boosting structured prediction for imitation learning. In Advances in Neural Information Processing Systems . 1153–1160
Nathan Ratliff, David M Bradley, Joel Chestnutt, and J A Bagnell. 2007 · 2007
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 3 (Chicago, Illinois) (AAAI’08) . AAAI Press, 1433–1438
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey. 2008 · 2008
Earlier work this paper cites.
Feature construction for inverse reinforcement learning. In Advances in Neural Information Processing Systems . 1342–1350
Sergey Levine, Zoran Popovic, and Vladlen Koltun. 2010 · 2010
Earlier work this paper cites.
Inverse reinforcement learning in partially observable environments
Jaedeug Choi and Kee-Eung Kim. 2011 · 2011
Cited alongside, same era.
Nonlinear inverse reinforcement learning with gaussian processes. In Advances in Neural Information Processing Systems . 19–27
Sergey Levine, Zoran Popovic, and Vladlen Koltun. 2011 · 2011
Cited alongside, same era.
Bayesian nonparametric feature construction for inverse reinforcement learning. In Twenty-Third International Joint Conference on Artificial Intelligence
Jaedeug Choi and Kee-Eung Kim. 2013 · 2013
Cited alongside, same era.
Finding Locally Optimal, Collision-Free Trajectories with Sequential Convex Optimization.. In Robotics: science and systems , Vol. 9. Citeseer, 1–10
John Schulman, Jonathan Ho, Alex X Lee, Ibrahim Awwal, Henry Bradlow, and Pieter Abbeel. 2013 · 2013
Cited alongside, same era.
Learning preferences for manipulation tasks from online coactive feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena. 2015 · 2015
Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In International Conference on Learning Representations
Justin Fu, Katie Luo, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Teaching inverse reinforcement learners via features and demonstrations. In Advances in Neural Information Processing Systems . 8464–8473
Luis Haug, Sebastian Tschiatschek, and Adish Singla. 2018 · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in Atari. In Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc., 8011–8023
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. 2018 · 2018
Later among the works it cites.
Martín Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 (New York, NY, USA) (ICML’16) . JMLR.org, 49–58
Chelsea Finn, Sergey Levine, and Pieter Abbeel. 2016 · 2016
Cited alongside, same era.
Learning Robot Objectives from Physical Human Interaction. In CoRL
Andrea Bajcsy, Dylan P. Losey, Marcia Kilchenman O’Malley, and Anca D. Dragan. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Cited alongside, same era.
Learning from Physical Human Corrections, One Feature at a Time. In Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction (Chicago, IL, USA) (HRI ’18) . ACM, New York, NY, USA, 141–149
Andrea Bajcsy, Dylan P. Losey, Marcia K. O’Malley, and Anca D. Dragan. 2018 · 2018
Cited alongside, same era.
Probabilistically Safe Robot Planning with Confidence-Based Human Predictions
Jaime F Fisac, Andrea Bajcsy, Sylvia L Herbert, David Fridovich-Keil, Steven Wang, Claire J Tomlin, and Anca D Dragan. 2018 · 2018
Cited alongside, same era.
Finding Locally Optimal, Collision-Free Trajectories with Sequential Convex Optimization
John Schulman, Jonathan Ho, Alex Lee, Ibrahim Awwal, Henry Bradlow, and Pieter Abbeel. [n.d.]
Cited in the paper.
PyBullet, a Python module for physics simulation for games, robotics and machine learning
Erwin Coumans and Yunfei Bai. 2016–2019 · 2019
Later among the works it cites.
Confidence-aware motion prediction for real-time collision avoidance
David Fridovich-Keil, Andrea Bajcsy, Jaime F. Fisac, Sylvia L. Herbert, Steven Wang, Anca D. Dragan, and Claire J. Tomlin. 2019 · 2019
Later among the works it cites.
Quantifying Hypothesis Space Misspecification in Learning From Human–Robot Demonstrations and Physical Corrections
A. Bobu, A. Bajcsy, J. F. Fisac, S. Deglurkar, and A. D. Dragan. 2020 · 2020
Closest in time.
SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net
Siddharth Reddy, Anca D. Dragan, and Sergey Levine. 2020 · 2020
Closest in time.
Watch this: Scalable cost-function learning for path planning in urban environments. In 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . 2089–2095
M. Wulfmeier, D. Z. Wang, and I. Posner. 2016 · 2095
Closest in time.