Fetching the paper…
Reading the bibliography…
Robots can learn the right reward function by querying a human expert.
Incorporating thresholds of indifference in probabilistic choice models
K. Krishnan · 1977
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
G. L. Nemhauser, L. A. Wolsey, and M. L. Fisher · 1978
Earlier work this paper cites.
Discrete choice analysis: theory and application to travel demand , volume 9
M. E. Ben-Akiva, S. R. Lerman, and S. R. Lerman · 1985
Earlier work this paper cites.
Information-based objective functions for active data selection
D. J. MacKay · 1992
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Fast sparse gaussian process methods: The informative vector machine
R. Herbrich, N. D. Lawrence, and M. Seeger · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
P. Abbeel and A. Y. Ng · 2005
Earlier work this paper cites.
Efficient test selection in active diagnosis via entropy approximation
A. X. Zheng, I. Rish, and A. Beygelzimer · 2005
Earlier work this paper cites.
On semi-supervised classification
B. Krishnapuram, D. Williams, Y. Xue, L. Carin, M. Figueiredo, and A. J. Hartemink · 2005
Earlier work this paper cites.
Cortical substrates for exploratory decisions in humans
N. D. Daw, J. P. O’doherty, P. Dayan, B. Seymour, and R. J. Dolan · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Cost-minimising strategies for data labelling: optimal stopping and active learning
C. Dimitrakakis and C. Savu-Krohn · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Active learning literature survey
B. Settles · 2009
Earlier work this paper cites.
A rational model of preference learning and choice prediction by children
C. G. Lucas, T. L. Griffiths, F. Xu, and C. Fawcett · 2009
Earlier work this paper cites.
Near-optimal bayesian active learning with noisy observations
D. Golovin, A. Krause, and D. Ray · 2010
Cited alongside, same era.
Real-time multiattribute bayesian preference elicitation with pairwise comparison queries
S. Guo and S. Sanner · 2010
Cited alongside, same era.
Optimal bayesian recommendation sets and myopically optimal choice query sets
P. Viappiani and C. Boutilier · 2010
Cited alongside, same era.
Adaptive submodularity: Theory and applications in active learning and stochastic optimization
D. Golovin and A. Krause · 2011
Cited alongside, same era.
Bayesian active learning for classification and preference learning
N. Houlsby, F. Huszár, Z. Ghahramani, and M. Lengyel · 2011
Cited alongside, same era.
Simultaneous learning and covering with adversarial noise
Active comparison based learning incorporating user uncertainty and noise
R. Holladay, S. Javdani, A. Dragan, and S. Srinivasa · 2016
Later among the works it cites.
Planning for autonomous cars that leverage effects on human actions
D. Sadigh, S. Sastry, S. A. Seshia, and A. D. Dragan · 2016
Later among the works it cites.
Fetch and freight: Standard platforms for service robot applications
M. Wise, M. Ferguson, D. King, E. Diehr, and D. Dymesich · 2016
Later among the works it cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Later among the works it cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. S. Sastry, and S. A. Seshia · 2017
Later among the works it cites.
Deep reinforcement learning from human preferences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Guillory and J. A. Bilmes · 2011
Cited alongside, same era.
Keyframe-based learning from demonstration
B. Akgun, M. Cakmak, K. Jiang, and A. L. Thomaz · 2012
Cited alongside, same era.
April: Active preference learning-based reinforcement learning
R. Akrour, M. Schoenauer, and M. Sebag · 2012
Cited alongside, same era.
Designing robot learners that ask good questions
M. Cakmak and A. L. Thomaz · 2012
Cited alongside, same era.
An active learning algorithm for ranking from pairwise preferences with an almost optimal query complexity
N. Ailon · 2012
Cited alongside, same era.
Elements of information theory
T. M. Cover and J. A. Thomas · 2012
Cited alongside, same era.
Individual choice behavior: A theoretical analysis
R. D. Luce · 2012
Cited alongside, same era.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Later among the works it cites.
Active reward learning from critiques
Y. Cui and S. Niekum · 2018
Later among the works it cites.
Learning from physical human corrections, one feature at a time
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2018
Later among the works it cites.
Batch active preference-based learning of reward functions
E. Biyik and D. Sadigh · 2018
Later among the works it cites.
Learning from richer human guidance: Augmenting comparison-based learning with feature queries
C. Basu, M. Singhal, and A. D. Dragan · 2018
Later among the works it cites.
Learning reward functions by integrating human demonstrations and preferences
M. Palan, G. Shevchuk, N. C. Landolfi, and D. Sadigh · 2019
Closest in time.
The green choice: Learning and influencing human decisions on shared roads
E. Biyik, D. A. Lazar, D. Sadigh, and R. Pedarsani · 2019
Closest in time.
Learning an urban air mobility encounter model from expert preferences
S. M. Katz, A.-C. L. Bihan, and M. J. Kochenderfer · 2019
Closest in time.
Active learning of reward dynamics from hierarchical queries
C. Basu, E. Biyik, Z. He, M. Singhal, and D. Sadigh · 2019
Closest in time.
Teacher-aware active robot learning
M. Racca, A. Oulasvirta, and V. Kyrki · 2019
Closest in time.
Batch active learning using determinantal point processes
E. Bıyık, K. Wang, N. Anari, and D. Sadigh · 2019
Closest in time.