Fetching the paper…
Reading the bibliography…
Today's robots are increasingly interacting with people and need to efficiently learn inexperienced user's preferences.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Gaussian processes for ordinal regression
W. Chu, Z. Ghahramani, and C. K. Williams · 2005
Earlier work this paper cites.
Learning by demonstration with critique from a human teacher
B. Argall, B. Browning, and M. Veloso · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
All of Statistics: A Concise Course in Statistical Inference
L. Wasserman · 2010
Earlier work this paper cites.
Human preferences for robot-human hand-over configurations
M. Cakmak, S. S. Srinivasa, M. K. Lee, J. Forlizzi, and S. Kiesler · 2011
Earlier work this paper cites.
Trajectories and keyframes for kinesthetic teaching: A human-robot interaction perspective
B. Akgun, M. Cakmak, J. W. Yoo, and A. L. Thomaz · 2012
Earlier work this paper cites.
Keyframe-based learning from demonstration
B. Akgun, M. Cakmak, K. Jiang, and A. L. Thomaz · 2012
Earlier work this paper cites.
Active comparison based learning incorporating user uncertainty and noise
R. Holladay, S. Javdani, A. Dragan, and S. Srinivasa · 2016
Earlier work this paper cites.
Fetch and freight: Standard platforms for service robot applications
M. Wise, M. Ferguson, D. King, E. Diehr, and D. Dymesich · 2016
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
Learning robot objectives from physical human interaction
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
C. Wirth, R. Akrour, G. Neumann, J. Fürnkranz, et al · 2017
Earlier work this paper cites.
Learning from richer human guidance: Augmenting comparison-based learning with feature queries
C. Basu, M. Singhal, and A. D. Dragan · 2018
Cited alongside, same era.
Modeling driver behavior from demonstrations in dynamic environments using spatiotemporal lattices
D. S. González, O. Erkent, V. Romero-Cano, J. Dibangoye, and C. Laugier · 2018
Cited alongside, same era.
Including uncertainty when learning from human corrections
D. P. Losey and M. K. O’Malley · 2018
Cited alongside, same era.
Active reward learning from critiques
Y. Cui and S. Niekum · 2018
Cited alongside, same era.
Survey on human–robot collaboration in industrial settings: Safety, intuitive interfaces and applications
V. Villani, F. Pini, F. Leali, and C. Secchi · 2018
Cited alongside, same era.
Asking easy questions: A user-friendly approach to active reward learning
Active preference learning using maximum regret
N. Wilde, D. Kulić, and S. L. Smith · 2020
Later among the works it cites.
Interactive robot training for non-markov tasks
A. Shah, S. Wadhwania, and J. Shah · 2020
Later among the works it cites.
Improving user specifications for robot behavior through active preference learning: Framework and evaluation
N. Wilde, A. Blidaru, S. L. Smith, and D. Kulić · 2020
Later among the works it cites.
Learning human-aware robot navigation from physical interaction via inverse reinforcement learning
M. Kollmitz, T. Koller, J. Boedecker, and W. Burgard · 2020
Later among the works it cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
D. S. Brown, W. Goo, and S. Niekum · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Biyik, M. Palan, N. C. Landolfi, D. P. Losey, and D. Sadigh · 2019
Cited alongside, same era.
Bayesian active learning for collaborative task specification using equivalence regions
N. Wilde, D. Kulić, and S. L. Smith · 2019
Cited alongside, same era.
Learning reward functions by integrating human demonstrations and preferences
M. Palan, N. C. Landolfi, G. Shevchuk, and D. Sadigh · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
D. Brown, W. Goo, P. Nagarajan, and S. Niekum · 2019
Cited alongside, same era.
Learning from extrapolated corrections
J. Y. Zhang and A. D. Dragan · 2019
Cited alongside, same era.
Reward-rational (implicit) choice: A unifying formalism for reward learning
H. J. Jeon, S. Milli, and A. D. Dragan · 2020
Cited alongside, same era.
Active preference-based gaussian process regression for reward learning
E. Biyik, N. Huynh, M. J. Kochenderfer, and D. Sadigh · 2020
Cited alongside, same era.
Learning from suboptimal demonstration via self-supervised reward regression
L. Chen, R. Paleja, and M. Gombolay · 2020
Later among the works it cites.
Controlling assistive robots with learned latent actions
D. P. Losey, K. Srinivasan, A. Mandlekar, A. Garg, and D. Sadigh · 2020
Later among the works it cites.
Interactive tuning of robot program parameters via expected divergence maximization
M. Racca, V. Kyrki, and M. Cakmak · 2020
Later among the works it cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
S. Cabi, S. G. Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, et al · 2020
Later among the works it cites.
Roial: Region of interest active learning for characterizing exoskeleton gait preference landscapes
K. Li, M. Tucker, E. Biyik, E. Novoseller, J. W. Burdick, Y. Sui, D. Sadigh, Y. Yue, and A. D. Ames · 2021
Closest in time.
Learning human objectives from sequences of physical corrections
M. Li, A. Canberk, D. P. Losey, and D. Sadigh · 2021
Closest in time.
Learning multimodal rewards from rankings
V. Myers, E. Biyik, N. Anari, and D. Sadigh · 2021
Closest in time.