Fetching the paper…
Reading the bibliography…
Effective robot learning often requires online human feedback and interventions that can cost significant human time, giving rise to the central challenge in interactive imitation learning: is it possible to control the timing and length of interventions to both facilitate learning and limit burden on the human supervisor? This paper presents ThriftyDAgger, an algorithm for actively querying a human supervisor given a desired budget of human interventions.
An adaptive hysteresis-band current control technique of a voltage-fed PWM inverter for machine drive system
B. K. Bose · 1990
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
D. A. Pomerleau · 1991
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Validating human-robot interaction schemes in multitasking environments
J. W. Crandall, M. A. Goodrich, D. R. Olsen, and C. W. Nielsen · 2005
Earlier work this paper cites.
NASA-task load index (NASA-TLX); 20 years later
S. G. Hart · 2006
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Supervisory control of multiple robots: Human-performance issues and user-interface design
J. Y. Chen, M. J. Barnes, and M. Harper-Sciarini · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and J. A. Bagnell · 2011
Earlier work this paper cites.
Active imitation learning via state queries
K. Judah, A. Fern, and T. Dietterich · 2011
Earlier work this paper cites.
MuJoCo: A Physics Engine for Model-Based Control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Dynamical movement primitives: learning attractor models for motor behaviors
A. J. Ijspeert, J. Nakanishi, H. Hoffmann, P. Pastor, and S. Schaal · 2013
Earlier work this paper cites.
Probabilistic movement primitives
A. Paraschos, C. Daniel, J. R. Peters, and G. Neumann · 2013
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning, 2014
S. Ross and J. A. Bagnell · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Earlier work this paper cites.
An open-source research kit for the da Vinci® surgical system
P. Kazanzides, Z. Chen, A. Deguet, G. S. Fischer, R. H. Taylor, and S. P. DiMaio · 2014
Earlier work this paper cites.
An invitation to imitation
J. A. Bagnell · 2015
Earlier work this paper cites.
SHIV: Reducing supervisor burden using support vectors for efficient learning from demonstrations in high dimensional state spaces
M. Laskey, S. Staszak, W. Hsieh, J. Mahler, F. Pokorny, A. Dragan, and K. Goldberg · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
DART: Noise injection for robust imitation learning
M. Laskey, J. Lee, R. Fox, A. Dragan, and K. Goldberg · 2017
Cited alongside, same era.
Query-efficient imitation learning for end-to-end autonomous driving
J. Zhang and K. Cho · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. S. Sastry, and S. A. Seshia · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
EnsembleDAgger: A Bayesian Approach to Safe Imitation Learning
K. Menda, K. Driggs-Campbell, and M. J. Kochenderfer · 2019
Later among the works it cites.
Exploring the limitations of behavior cloning for autonomous driving
F. Codevilla, E. Santana, L. A. M., and A. Gaidon · 2019
Later among the works it cites.
Better-than-demonstrator imitation learning via automaticaly-ranked demonstrations
D. S. Brown, W. Goo, and S. Niekum · 2019
Later among the works it cites.
On-policy robot imitation learning from a converging supervisor
A. Balakrishna*, B. Thananjeyan*, J. Lee, F. Li, A. Zahed, J. E. Gonzalez, and K. Goldberg · 2019
Later among the works it cites.
Learning reward functions by integrating human demonstrations and preferences
M. Palan, N. C. Landolfi, G. Shevchuk, and D. Sadigh · 2019
Later among the works it cites.
Asking easy questions: A user-friendly approach to active reward learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, and J. Peters · 2018
Cited alongside, same era.
Agile autonomous driving using end-to-end deep imitation learning
Y. Pan, C.-A. Cheng, K. Saigol, K. Lee, X. Yan, E. Theodorou, and B. Boots · 2018
Cited alongside, same era.
End-to-end driving via conditional imitation learning
F. Codevilla, M. Müller, A. López, V. Koltun, and A. Dosovitskiy · 2018
Cited alongside, same era.
Deep reinforcement learning doesn’t work yet
A. Irpan · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine · 2018
Cited alongside, same era.
A survey of inverse reinforcement learning: Challenges, methods and progress
S. Arora and P. Doshi · 2018
Cited alongside, same era.
E. Bıyık, M. Palan, N. C. Landolfi, D. P. Losey, and D. Sadigh · 2019
Later among the works it cites.
Learning interpretable and transferable rope manipulation policies using depth sensing and dense object descriptors
P. Sundaresan, P. Grannen, B. Thananjeyan, A. Balakrishna, M. Laskey, K. Stone, J. E. Gonzalez, and K. Goldberg · 2020
Later among the works it cites.
Learning from interventions: Human-robot interaction as both explicit and implicit feedback
J. Spencer, S. Choudhury, M. Barnes, M. Schmittle, M. Chiang, P. Ramadge, and S. Srinivasa · 2020
Later among the works it cites.
Interactive imitation learning in state-space
S. Jauhri, C. Celemin, and J. Kober · 2020
Later among the works it cites.
Scaled autonomy: Enabling human operators to control robot fleets
G. Swamy, S. Reddy, S. Levine, and A. D. Dragan · 2020
Later among the works it cites.
Human-in-the-loop imitation learning using remote teleoperation, 2020
A. Mandlekar, D. Xu, R. Martín-Martín, Y. Zhu, L. Fei-Fei, and S. Savarese · 2020
Later among the works it cites.
Safe imitation learning via fast Bayesian reward inference from preferences
D. Brown, R. Coleman, R. Srinivasan, and S. Niekum · 2020
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
B. Thananjeyan*, A. Balakrishna*, S. Nair, M. Luo, K. Srinivasan, M. Hwang, and J. E. Gonzalez · 2020
Later among the works it cites.
Robosuite: A modular simulation framework and benchmark for robot learning
Y. Zhu, J. Wong, A. Mandlekar, and R. Martín-Martín · 2020
Later among the works it cites.
Learning Dense Visual Correspondences in Simulation to Smooth and Fold Real Fabrics
A. Ganapathi, P. Sundaresan, B. Thananjeyan, A. Balakrishna, D. Seita, J. Grannen, M. Hwang, R. Hoque, J. E. Gonzalez, N. Jamali, K. Yamane, S. Iba, and K. Goldberg · 2021
Closest in time.
A review of robot learning for manipulation: Challenges, representations, and algorithms
O. Kroemer, S. Niekum, and G. Konidaris · 2021
Closest in time.
LazyDAgger: Reducing context switching in interactive imitation learning
R. Hoque, A. Balakrishna, C. Putterman, M. Luo, D. S. Brown, D. Seita, B. Thananjeyan, E. Novoseller, and K. Goldberg · 2021
Closest in time.