Fetching the paper…
Reading the bibliography…
The goal of imitation learning is to mimic expert behavior without access to an explicit reward signal.
D. A. Pomerleau, “Alvinn, an autonomous land vehicle in a neural network,” tech. rep., Carnegie Mellon University, Computer Science Department, 1989
1989
Earlier work this paper cites.
D. A. Pomerleau, “Efficient training of artificial neural networks for autonomous navigation,” Neural Computation
1991
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning
1992
Earlier work this paper cites.
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner, “Torcs, the open racing car simulator,” Software available at http://torcs. sourceforge. net
2000
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning
2004
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.,” in AAAI
2008
Earlier work this paper cites.
U. Syed, M. Bowling, and R. E. Schapire, “Apprenticeship learning using linear programming,” in Proceedings of the 25th international conference on Machine learning
2008
Earlier work this paper cites.
M. Aly, “Real time detection of lane markers in urban streets,” in Intelligent Vehicles Symposium, 2008 IEEE
2008
Earlier work this paper cites.
S. Ross and D. Bagnell, “Efficient reductions for imitation learning.,” in AISTATS
2010
Earlier work this paper cites.
S. Ross, G. J. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning.,” in AISTATS
2011
Earlier work this paper cites.
P. Lenz, J. Ziegler, A. Geiger, and M. Roser, “Sparse scene flow segmentation for moving object detection in urban environments,” in Intelligent Vehicles Symposium (IV), 2011 IEEE
2011
Earlier work this paper cites.
K. Kitani, B. Ziebart, J. Bagnell, and M. Hebert, “Activity forecasting,” Computer Vision–ECCV 2012
2012
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search.,” in ICML (3)
2013
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,” The International Journal of Robotics Research
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems
2014
Cited alongside, same era.
M. Bloem and N. Bambos, “Infinite time horizon maximum causal entropy inverse reinforcement learning,” in Decision and Control (CDC), 2014 IEEE 53rd Annual Conference on
2014
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980
2014
Cited alongside, same era.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?,” in Advances in neural information processing systems
2014
Cited alongside, same era.
A. Tamar, S. Levine, P. Abbeel, Y. WU, and G. Thomas, “Value iteration networks,” in Advances in Neural Information Processing Systems
2016
Later among the works it cites.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in Proceedings of the 33rd International Conference on Machine Learning
2016
Later among the works it cites.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” in Advances in Neural Information Processing Systems
2016
Later among the works it cites.
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel, “Infogan: Interpretable representation learning by information maximizing generative adversarial nets,” in Advances in Neural Information Processing Systems
2016
Later among the works it cites.
J. Ho, J. K. Gupta, and S. Ermon, “Model-free imitation learning with policy optimization,” in Proceedings of the 33rd International Conference on Machine Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization.,” in ICML
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
P. Englert and M. Toussaint, “Inverse kkt–learning cost functions of manipulation tasks from demonstrations,” in Proceedings of the International Symposium of Robotics Research
2015
Cited alongside, same era.
S. Ermon, Y. Xue, R. Toth, B. N. Dilkina, R. Bernstein, T. Damoulas, P. Clark, S. DeGloria, A. Mude, C. Barrett, et al
2015
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Cited alongside, same era.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al
2015
Cited alongside, same era.
2016
Later among the works it cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” in Advances in Neural Information Processing Systems
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
B. Stadie, P. Abbeel, and I. Sutskever, “Third person imitation learning,” in ICLR
2017
Closest in time.
2017
Closest in time.