Fetching the paper…
Reading the bibliography…
Our goal is to accurately and efficiently learn reward functions for autonomous robots.
Conditional logit analysis of qualitative choice behavior
Daniel McFadden et al · 1973
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Maximum margin planning
Nathan D Ratliff, J Andrew Bagnell, and Martin A Zinkevich · 2006
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Active preference learning with discrete choice data
Brochu Eric, Nando D Freitas, and Abhijeet Ghosh · 2008
Earlier work this paper cites.
Exploring voting blocs within the irish electorate: A mixture modeling approach
Isobel Claire Gormley and Thomas Brendan Murphy · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Bayesian inference for plackett-luce ranking models
John Guiver and Edward Snelson · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, Jonathan Eckstein, et al · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Integrating reinforcement learning with human demonstrations of varying ability
Matthew E Taylor, Halit Bener Suay, and Sonia Chernova · 2011
Earlier work this paper cites.
An active learning algorithm for ranking from pairwise preferences with an almost optimal query complexity
Nir Ailon · 2012
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
Riad Akrour, Marc Schoenauer, and Michèle Sebag · 2012
Earlier work this paper cites.
Formalizing assistive teleoperation
Anca D Dragan and Siddhartha S Srinivasa · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park · 2012
Cited alongside, same era.
Continuous inverse optimal control with locally optimal examples
Sergey Levine and Vladlen Koltun · 2012
Cited alongside, same era.
Individual choice behavior: A theoretical analysis
R Duncan Luce · 2012
Cited alongside, same era.
Preference-learning based inverse reinforcement learning for dialog control
Hiroaki Sugiyama, Toyomi Meguro, and Yasuhiro Minami · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
A bayesian approach for policy learning from trajectory preference queries
Learning robot objectives from physical human interaction
Andrea Bajcsy, Dylan P Losey, Marcia K O’Malley, and Anca D Dragan · 2017
Later among the works it cites.
Do you want your autonomous car to drive like you?
Chandrayee Basu, Qian Yang, David Hungerman, Mukesh Singhal, and Anca D Dragan · 2017
Later among the works it cites.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Later among the works it cites.
One-shot imitation learning
Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Later among the works it cites.
URL https://twitter.com/mat_kelcey/status/886101319559335936
Mat Kelcey, 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Cited alongside, same era.
Generating legible motion
Anca Dragan and Siddhartha Srinivasa · 2013
Cited alongside, same era.
Learning preferences for manipulation tasks from online coactive feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena · 2015
Cited alongside, same era.
Shared autonomy via hindsight optimization
Shervin Javdani, Siddhartha S Srinivasa, and J Andrew Bagnell · 2015
Cited alongside, same era.
Data-driven motion mappings improve transparency in teleoperation
Rebecca P Khurshid and Katherine J Kuchenbecker · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Batch active preference-based learning of reward functions
Erdem Bıyık and Dorsa Sadigh · 2018
Later among the works it cites.
Make the table/big block in fetch environments fixed., 2018
Joy Chopra · 2018
Later among the works it cites.
Active reward learning from critiques
Yuchen Cui and Scott Niekum · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Later among the works it cites.
An efficient, generalized bellman update for cooperative inverse reinforcement learning
Dhruv Malik, Malayandi Palaniappan, Jaime F Fisac, Dylan Hadfield-Menell, Stuart Russell, and Anca D Dragan · 2018
Later among the works it cites.
Simplifying reward design through divide-and-conquer
Ellis Ratner, Dylan Hadfield-Menell, and Anca D Dragan · 2018
Later among the works it cites.
Risk-aware active inverse reinforcement learning
Daniel S Brown, Yuchen Cui, and Scott Niekum · 2019
Closest in time.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2019
Closest in time.