Fetching the paper…
Reading the bibliography…
In many settings, it is desirable to learn decision-making and control policies through learning or bootstrapping from expert demonstrations.
Divergence measures based on the shannon entropy
J. Lin · 1991
Earlier work this paper cites.
Robot learning from demonstration
C. G. Atkeson and S. Schaal · 1997
Earlier work this paper cites.
Learning from demonstration
S. Schaal · 1997
Earlier work this paper cites.
Learning agents for uncertain environments
S. J. Russell · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
E. Todorov · 2007
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Learning for control from multiple demonstrations
A. Coates, P. Abbeel, and A. Y. Ng · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
M. Toussaint · 2009
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mulling, and Y. Altun · 2010
Cited alongside, same era.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Cited alongside, same era.
Autonomous helicopter aerobatics through apprenticeship learning
P. Abbeel, A. Coates, and A. Y. Ng · 2010
Cited alongside, same era.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
X. Nguyen, M. J. Wainwright, and M. I. Jordan · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Cited alongside, same era.
Optimal control as a graphical model inference problem
H. J. Kappen, V. Gómez, and M. Opper · 2012
Cited alongside, same era.
f-gan: Training generative neural samplers using variational divergence minimization
S. Nowozin, B. Cseke, and R. Tomioka · 2016
Later among the works it cites.
Reward augmented maximum likelihood for neural structured prediction
M. Norouzi, S. Bengio, N. Jaitly, M. Schuster, Y. Wu, D. Schuurmans, et al · 2016
Later among the works it cites.
Deep learning , volume 1
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio · 2016
Later among the works it cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Later among the works it cites.
Iterative noise injection for scalable imitation learning
M. Laskey, J. Lee, W. Hsieh, R. Liaw, J. Mahler, R. Fox, and K. Goldberg · 2017
Later among the works it cites.
Improved training of wasserstein gans
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Infinite time horizon maximum causal entropy inverse reinforcement learning
M. Bloem and N. Bambos · 2014
Cited alongside, same era.
Training generative neural networks via maximum mean discrepancy optimization
G. K. Dziugaite, D. M. Roy, and Z. Ghahramani · 2015
Cited alongside, same era.
Generative moment matching networks
Y. Li, K. Swersky, and R. Zemel · 2015
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville · 2017
Later among the works it cites.
Learning robust rewards with adverserial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Later among the works it cites.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, et al · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Later among the works it cites.
Formal limitations on the measurement of mutual information
D. McAllester and K. Statos · 2018
Later among the works it cites.
Efficient exploration via state marginal matching
L. Lee, B. Eysenbach, E. Parisotto, E. Xing, S. Levine, and R. Salakhutdinov · 2019
Closest in time.
Imitation learning as f f -divergence minimization
L. Ke, M. Barnes, W. Sun, G. Lee, S. Choudhury, and S. Srinivasa · 2019
Closest in time.