Fetching the paper…
Reading the bibliography…
GAIL is a recent successful imitation learning architecture that exploits the adversarial training procedure introduced in GANs.
ALVINN: An Autonomous Land Vehicle in a Neural Network
Pomerleau, D. (1989) · 1989
Earlier work this paper cites.
Rapidly Adapting Artificial Neural Networks for Autonomous Navigation
Pomerleau, D. (1990) · 1990
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. (1999) · 1999
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
Apprenticeship Learning via Inverse Reinforcement Learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Robot Programming by Demonstration
Billard, A., Calinon, S., Dillmann, R., and Schaal, S. (2008) · 2008
Earlier work this paper cites.
Apprenticeship Learning Using Linear Programming
Syed, U., Bowling, M., and Schapire, R. E. (2008) · 2008
Earlier work this paper cites.
A Game-Theoretic Approach to Apprenticeship Learning
Syed, U. and Schapire, R. E. (2008) · 2008
Earlier work this paper cites.
Efficient Reductions for Imitation Learning
Ross, S. and Bagnell, J. A. (2010) · 2010
Earlier work this paper cites.
Nonlinear Inverse Reinforcement Learning with Gaussian Processes
Levine, S., Popovic, Z., and Koltun, V. (2011) · 2011
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Ross, S., Gordon, G. J., and Bagnell, J. A. (2011) · 2011
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Deterministic Policy Gradient Algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
An invitation to imitation
Bagnell, J. A. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
Layer Normalization
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Cited alongside, same era.
OpenAI Gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
Finn, C., Christiano, P., Abbeel, P., and Levine, S. (2016) · 2016
Noisy Networks for Exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S. (2017) · 2017
Later among the works it cites.
Improved Training of Wasserstein GANs
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. (2017) · 2017
Later among the works it cites.
Multi-Modal Imitation Learning from Unstructured Demonstrations using Generative Adversarial Nets
Hausman, K., Chebotar, Y., Schaal, S., Sukhatme, G., and Lim, J. (2017) · 2017
Later among the works it cites.
Deep Reinforcement Learning that Matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2017) · 2017
Later among the works it cites.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2017) · 2017
Later among the works it cites.
Categorical Reparameterization with Gumbel-Softmax
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
NIPS 2016 Tutorial: Generative Adversarial Networks
Goodfellow, I. (2017) · 2016
Cited alongside, same era.
Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S. (2016) · 2016
Cited alongside, same era.
Generative Adversarial Imitation Learning
Ho, J. and Ermon, S. (2016) · 2016
Cited alongside, same era.
Model-Free Imitation Learning with Policy Optimization
Ho, J., Gupta, J. K., and Ermon, S. (2016) · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Connecting Generative Adversarial Networks and Actor-Critic Methods
Pfau, D. and Vinyals, O. (2016) · 2016
Cited alongside, same era.
Jang, E., Gu, S., and Poole, B. (2017) · 2017
Later among the works it cites.
Burn-In Demonstrations for Multi-Modal Imitation Learning
Kuefler, A. and Kochenderfer, M. J. (2017) · 2017
Later among the works it cites.
InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
Li, Y., Song, J., and Ermon, S. (2017) · 2017
Later among the works it cites.
Are GANs Created Equal? A Large-Scale Study
Lucic, M., Kurach, K., Michalski, M., Gelly, S., and Bousquet, O. (2017) · 2017
Later among the works it cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Maddison, C. J., Mnih, A., and Teh, Y. W. (2017) · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Oleg, K. (2017) · 2017
Later among the works it cites.
Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Večerík, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M. (2017) · 2017
Later among the works it cites.
Energy-based Generative Adversarial Network
Zhao, J., Mathieu, M., and LeCun, Y. (2017) · 2017
Later among the works it cites.
Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
Fu, J., Luo, K., and Levine, S. (2018) · 2018
Closest in time.
Addressing Sample Inefficiency and Reward Bias in Inverse Reinforcement Learning
Kostrikov, I., Agrawal, K. K., Levine, S., and Tompson, J. (2018) · 2018
Closest in time.
Parameter Space Noise for Exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M. (2018) · 2018
Closest in time.
Deterministic Policy Imitation Gradient Algorithm
Sasaki, F. and Kawaguchi, A. (2018) · 2018
Closest in time.