Fetching the paper…
Reading the bibliography…
While imitation learning is becoming common practice in robotics, this approach often suffers from data mismatch and compounding errors.
S. Schaal, “Learning from demonstration,” in Advances in Neural Information Processing Systems (NIPS) , 1997, pp. 1040–1046
1997
Earlier work this paper cites.
B. Price and C. Boutilier, “Accelerating reinforcement learning through implicit imitation,” Journal of Artificial Intelligence Research , vol. 19, pp. 569–629, 2003
2003
Earlier work this paper cites.
B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and Autonomous Systems , vol. 57, no. 5, pp. 469–483, 2009
2009
Earlier work this paper cites.
H. Daumé, J. Langford, and D. Marcu, “Search-based structured prediction,” Machine Learning , vol. 75, no. 3, pp. 297–325, 2009
2009
Earlier work this paper cites.
J. Kober and J. Peters, “Imitation and reinforcement learning,” IEEE Robotics Automation Magazine , vol. 17, no. 2, pp. 55–62, June 2010
2010
Earlier work this paper cites.
S. Ross and D. Bagnell, “Efficient reductions for imitation learning,” in International Conference on Artificial Intelligence and Statistics , 2010, pp. 661–668
2010
Earlier work this paper cites.
S. Ross, G. J. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in International Conference on Artificial Intelligence and Statistics , 2011, pp. 627–635
2011
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning (ICML) , 2013, pp. 1–9
2013
Cited alongside, same era.
B. Kim and J. Pineau, “Maximum mean discrepancy imitation learning,” in Robotics: Science and Systems , 2013
2013
Cited alongside, same era.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” Journal of Machine Learning , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Cited alongside, same era.
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International Conference on Machine Learning (ICML) , 2015, pp. 1889–1897
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in International Conference on Machine Learning (ICML) , 2016, pp. 1329–1338
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2016
Cited alongside, same era.
M. Laskey, S. Staszak, W. Y.-S. Hsieh, J. Mahler, et al. , “Shiv: Reducing supervisor burden in dagger using support vectors for efficient learning from demonstrations in high dimensional state spaces,” in IEEE International Conference on Robotics and Automation (ICRA) , 2016, pp. 462–469
2016
Cited alongside, same era.
2017
Closest in time.