Fetching the paper…
Reading the bibliography…
Imitation learning (IL) consists of a set of tools that leverage expert demonstrations to quickly learn policies.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Bregman, L. M. (1967) · 1967
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A. (1989) · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Schaal, S. (1999) · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
On choosing and bounding probability metrics
Gibbs, A. L. and Su, F. E. (2002) · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J. (2004) · 2004
Earlier work this paper cites.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A. (2009) · 2009
Cited alongside, same era.
Relative entropy policy search
Peters, J., Mülling, K., and Altun, Y. (2010) · 2010
Cited alongside, same era.
First order methods for nonsmooth convex large-scale optimization, i: general purpose methods
Juditsky, A., Nemirovski, A., et al. (2011) · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, K., Toussaint, M., and Vijayakumar, S. (2012) · 2012
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y. (2017) · 2017
Later among the works it cites.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Grathwohl, W., Choi, D., Wu, Y., Roeder, G., and Duvenaud, D. (2017) · 2017
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Agile off-road autonomous driving using end-to-end deep imitation learning
Pan, Y., Cheng, C.-A., Saigol, K., Lee, K., Yan, X., Theodorou, E., and Boots, B. (2017) · 2017
Later among the works it cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Cited alongside, same era.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A. (2014) · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Learning to search better than your teacher
Chang, K.-W., Krishnamurthy, A., Agarwal, A., Daume III, H., and Langford, J. (2015) · 2015
Cited alongside, same era.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Ghadimi, S., Lan, G., and Zhang, H. (2016) · 2016
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015a)
Cited in the paper.
Rajeswaran, A., Kumar, V., Gupta, A., Schulman, J., Todorov, E., and Levine, S. (2017) · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017) · 2017
Later among the works it cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A. (2017) · 2017
Later among the works it cites.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, G., Mnih, A., Maddison, C. J., Lawson, J., and Sohl-Dickstein, J. (2017) · 2017
Later among the works it cites.
Convergence of value aggregation for imitation learning
Cheng, C.-A. and Boots, B. (2018) · 2018
Closest in time.
DART: Dynamic animation and robotics toolkit
Lee, J., Grey, M. X., Ha, S., Kunz, T., Jain, S., Ye, Y., Srinivasa, S. S., Stilman, M., and Liu, C. K. (2018) · 2018
Closest in time.
Truncated horizon policy search: Deep combination of reinforcement and imitation
Sun, W., Bagnell, J. A., and Boots, B. (2018) · 2018
Closest in time.