Fetching the paper…
Reading the bibliography…
It is often difficult to hand-specify what the correct reward function is for a task, so researchers have instead aimed to learn reward functions from human behavior or feedback.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Rational choice and the structure of the environment
Herbert A Simon · 1956
Earlier work this paper cites.
Information theory and statistical mechanics
Edwin T Jaynes · 1957
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R Duncan Luce · 1959
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Impedance control: An approach to manipulation: Part ii—implementation
Neville Hogan · 1985
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Maximum margin planning
Nathan D Ratliff, J Andrew Bagnell, and Martin A Zinkevich · 2006
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Collision detection and reaction: A contribution to safe physical human-robot interaction
Sami Haddadin, Alin Albu-Schaffer, Alessandro De Luca, and Gerd Hirzinger · 2008
Earlier work this paper cites.
Action understanding as inverse planning
Chris L Baker, Rebecca Saxe, and Joshua B Tenenbaum · 2009
Earlier work this paper cites.
Cause and intent: Social reasoning in causal learning
Noah D Goodman, Chris L Baker, and Joshua B Tenenbaum · 2009
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew R Walter, Ashis Gopal Banerjee, Seth Teller, and Nicholas Roy · 2011
Earlier work this paper cites.
A joint model of language and perception for grounded attribute learning
Cynthia Matuszek, Nicholas FitzGerald, Luke Zettlemoyer, Liefeng Bo, and Dieter Fox · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Cited alongside, same era.
Continuous inverse optimal control with locally optimal examples
Sergey Levine and Vladlen Koltun · 2012
Cited alongside, same era.
Formalizing assistive teleoperation
Anca D Dragan and Siddhartha S Srinivasa · 2012
Cited alongside, same era.
Knowledge and implicature: Modeling language understanding as social cognition
Noah D Goodman and Andreas Stuhlmüller · 2013
Cited alongside, same era.
Legibility and predictability of robot motion
Anca D Dragan, Kenton CT Lee, and Siddhartha S Srinivasa · 2013
Cited alongside, same era.
Policy shaping: Integrating human feedback with reinforcement learning
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz · 2013
Active comparison based learning incorporating user uncertainty and noise
Rachel Holladay, Shervin Javdani, Anca Dragan, and Siddhartha Srinivasa · 2016
Later among the works it cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Later among the works it cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Later among the works it cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D. Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Later among the works it cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Later among the works it cites.
Learning robot objectives from physical human interaction
Andrea Bajcsy, Dylan P Losey, Marcia K O’Malley, and Anca D Dragan · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Preference-based evolutionary direct policy search
Róbert Busa-Fekete, Balázs Szörényi, Paul Weng, Weiwei Cheng, and Eyke Hüllermeier · 2013
Cited alongside, same era.
Thermodynamics as a theory of decision-making with information-processing costs
Pedro A Ortega and Daniel A Braun · 2013
Cited alongside, same era.
Infinite time horizon maximum causal entropy inverse reinforcement learning
Michael Bloem and Nicholas Bambos · 2014
Cited alongside, same era.
On learning from game annotations
Christian Wirth and Johannes Fürnkranz · 2014
Cited alongside, same era.
A strategy-aware technique for learning behaviors from discrete human feedback
Robert Tyler Loftin, James MacGlashan, Bei Peng, Matthew E Taylor, Michael L Littman, Jeff Huang, and David L Roberts · 2014
Cited alongside, same era.
Motion planning with sequential convex optimization and convex collision checking
John Schulman, Yan Duan, Jonathan Ho, Alex Lee, Ibrahim Awwal, Henry Bradlow, Jia Pan, Sachin Patil, Ken Goldberg, and Pieter Abbeel · 2014
Cited alongside, same era.
Later among the works it cites.
Variational inference: A review for statisticians
David M Blei, Alp Kucukelbir, and Jon D McAuliffe · 2017
Later among the works it cites.
Trajectory deformations from physical human–robot interaction
Dylan P Losey and Marcia K O’Malley · 2017
Later among the works it cites.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman · 2017
Later among the works it cites.
Simplifying reward design through divide-and-conquer
Ellis Ratner, Dylan Hadfield-Menell, and Anca D Dragan · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Later among the works it cites.
Including uncertainty when learning from human corrections
Dylan P Losey and Marcia K O’Malley · 2018
Later among the works it cites.
Active inverse reward design
Sören Mindermann, Rohin Shah, Adam Gleave, and Dylan Hadfield-Menell · 2018
Later among the works it cites.
From language to goals: Inverse reinforcement learning for vision-based instruction following
Justin Fu, Anoop Korattikara, Sergey Levine, and Sergio Guadarrama · 2019
Later among the works it cites.
Preferences implicit in the state of the world
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Later among the works it cites.
Learning reward functions by integrating human demonstrations and preferences
Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh · 2019
Later among the works it cites.