Fetching the paper…
Reading the bibliography…
Learning from Demonstration (LfD) seeks to democratize robotics by enabling non-roboticist end-users to teach robots to perform a task by providing a human demonstration.
Human problem solving , volume 104
A. Newell and H. A. Simon · 1972
Earlier work this paper cites.
Obtaining good performance from a bad teacher
M. Kaiser, H. Friedrich, and R. Dillmann · 1995
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
The kuka-dlr lightweight robot arm - a new reference platform for robotics research and manufacturing
R. Bischoff, J. Kurth, G. Schreiber, R. Koeppe, A. Albu-Schaeffer, A. Beyer, O. Eiberger, S. Haddadin, A. Stemmer, G. Grunwald, and G. Hirzinger · 2010
Earlier work this paper cites.
Towards robot scientists for autonomous scientific discovery
A. Sparkes, W. Aubrey, E. Byrne, A. Clare, M. N. Khan, M. Liakata, M. Markham, J. Rowland, L. Soldatova, K. E. Whelan, M. Young, and R. King · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Learning table tennis with a mixture of motor primitives
K. Müelling, J. Kober, and J. Peters · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Preference-based policy learning
R. Akrour, M. Schoenauer, and M. Sebag · 2011
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
H. Sugiyama, T. Meguro, and Y. Minami · 2012
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R. D. Luce · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Towards a comprehensive chore list for domestic robots
M. Cakmak and L. Takayama · 2013
Cited alongside, same era.
Semi-supervised apprenticeship learning
M. Valko, M. Ghavamzadeh, and A. Lazaric · 2013
Cited alongside, same era.
Towards robot skill learning: From simple skills to table tennis
J. Peters, J. Kober, K. Mülling, O. Krämer, and G. Neumann · 2013
Cited alongside, same era.
Learning to select and generalize striking movements in robot table tennis
K. Mülling, J. Kober, O. Kroemer, and J. Peters · 2013
Cited alongside, same era.
Learning strategies in table tennis using inverse reinforcement learning
K. Muelling, A. Boularias, B. Mohler, B. Schölkopf, and J. Peters · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Learning robust rewards with adverserial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Later among the works it cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Online optimal trajectory generation for robot table tennis
O. Koç, G. Maeda, and J. Peters · 2018
Later among the works it cites.
A. Hill, A. Raffin, M. Ernestus, A. Gleave, A. Kanervisto, R. Traore, P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distance minimization for reward learning from scored trajectories
B. Burchfiel, C. Tomasi, and R. Parr · 2016
Cited alongside, same era.
Apprenticeship scheduling: Learning to schedule from human experts
M. Gombolay, R. Jensen, J. Stigile, S.-H. Son, and J. Shah · 2016
Cited alongside, same era.
Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning
S. Krishnan, A. Garg, S. Patil, C. Lea, G. Hager, P. Abbeel, and K. Goldberg · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Dart: Noise injection for robust imitation learning
M. Laskey, J. Lee, R. Fox, A. D. Dragan, and K. Goldberg · 2017
Cited alongside, same era.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
D. S. Brown, W. Goo, and S. Niekum · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
D. Brown, W. Goo, P. Nagarajan, and S. Niekum · 2019
Later among the works it cites.
Imitation learning from imperfect demonstration
Y.-H. Wu, N. Charoenphakdee, H. Bao, V. Tangkaratt, and M. Sugiyama · 2019
Later among the works it cites.
Heterogeneous graph attention networks for scalable multi-robot scheduling with temporospatial constraints
Z. Wang and M. Gombolay · 2020
Closest in time.
Recent advances in robot learning from demonstration
H. Ravichandar, A. S. Polydoros, S. Chernova, and A. Billard · 2020
Closest in time.
Joint goal and strategy inference across heterogeneous demonstrators via reward network distillation
L. Chen, R. R. Paleja, M. Ghuy, and M. C. Gombolay · 2020
Closest in time.
Less is more: Rethinking probabilistic models of human behavior
A. Bobu, D. R. Scobee, J. F. Fisac, S. S. Sastry, and A. D. Dragan · 2020
Closest in time.
Interpretable and personalized apprenticeship scheduling: Learning interpretable scheduling policies from heterogeneous user demonstrations
R. R. Paleja, A. Silva, L. Chen, and M. Gombolay · 2020
Closest in time.