Fetching the paper…
Reading the bibliography…
Imitation learning is well-suited for robotic tasks where it is difficult to directly program the behavior or specify a cost for optimal control.
A general class of coefficients of divergence of one distribution from another
S. M. Ali and S. D. Silvey · 1966
Earlier work this paper cites.
Sample estimate of the entropy of a random vector
L. Kozachenko and N. N. Leonenko · 1987
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1989
Earlier work this paper cites.
A framework for behavioural cloning
M. Bain and C. Sammut · 1995
Earlier work this paper cites.
Robot learning from demonstration
C. G. Atkeson and S. Schaal · 1997
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Müller · 1997
Earlier work this paper cites.
Learning agents for uncertain environments
S. Russell · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Non-parametric entropy estimation toolbox (npeet)
G. Ver Steeg · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Estimating mutual information
A. Kraskov, H. Stögbauer, and P. Grassberger · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Optimal transport: old and new , volume 338
C. Villani · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Maximum entropy deep inverse reinforcement learning
M. Wulfmeier, P. Ondruska, and I. Posner · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan · 2016
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Later among the works it cites.
Safe exploration in continuous action spaces
G. Dalal, K. Dvijotham, M. Vecerik, T. Hester, C. Paduraru, and Y. Tassa · 2018
Later among the works it cites.
Adversarial imitation via variational inverse reinforcement learning
A. H. Qureshi, B. Boots, and M. C. Yip · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
f-gan: Training generative neural samplers using variational divergence minimization
S. Nowozin, B. Cseke, and R. Tomioka · 2016
Cited alongside, same era.
Analysis of k-nearest neighbor distances with application to entropy estimation
S. Singh and B. Póczos · 2016
Cited alongside, same era.
Online trajectory planning and force control for automation of surgical tasks
T. Osa, N. Sugita, and M. Mitsuishi · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2017
Cited alongside, same era.
Later among the works it cites.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
S. K. S. Ghasemipour, R. Zemel, and S. Gu · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
L. Lee, B. Eysenbach, E. Parisotto, E. Xing, S. Levine, and R. Salakhutdinov · 2019
Later among the works it cites.
Random expert distillation: Imitation learning via expert policy support estimation
R. Wang, C. Ciliberto, P. Amadori, and Y. Demiris · 2019
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
S. Reddy, A. D. Dragan, and S. Levine · 2019
Later among the works it cites.
Disagreement-regularized imitation learning
K. Brantley, W. Sun, and M. Henaff · 2019
Later among the works it cites.
State alignment-based imitation learning
F. Liu, Z. Ling, T. Mu, and H. Su · 2019
Later among the works it cites.
Imitation learning as f f -divergence minimization
L. Ke, M. Barnes, W. Sun, G. Lee, S. Choudhury, and S. Srinivasa · 2019
Later among the works it cites.
Energy-based imitation learning
M. Liu, T. He, M. Xu, and W. Zhang · 2020
Closest in time.