Fetching the paper…
Reading the bibliography…
Imitation learning (IL) aims to learn an optimal policy from demonstrations.
On the method of bounded differences
Colin McDiarmid · 1989
Earlier work this paper cites.
A 1.8 å resolution potential function for protein folding
Gordon M Crippen and Mark E Snow · 1990
Earlier work this paper cites.
Probability in Banach Spaces: Isoperimetry and Processes
Michel Ledoux and Michel Talagrand · 1991
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Introduction to Reinforcement Learning , volume 135
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Statistical Learning Theory , volume 3
Vladimir Vapnik · 1998
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal · 1999
Earlier work this paper cites.
Advances in Kernel Methods: Support Vector Learning
Bernhard Schölkopf, Christopher J. C. Burges, and Alexander J. Smola · 1999
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Loss functions for binary class probability estimation and classification: Structure and applications
Andreas Buja, Werner Stuetzle, and Yi Shen · 2005
Earlier work this paper cites.
Semi-Supervised Learning
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien · 2006
Earlier work this paper cites.
Sparseness vs estimating conditional probabilities: Some asymptotic results
Peter L Bartlett and Ambuj Tewari · 2007
Earlier work this paper cites.
Apprenticeship learning using linear programming
Umar Syed, Michael Bowling, and Robert E Schapire · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
A semi-supervised learning approach for soft labeled data
Mohamed M El-Zahhar and Neamat F El-Gayar · 2010
Cited alongside, same era.
Composite binary losses
Mark D Reid and Robert C Williamson · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Semi-supervised apprenticeship learning
Michal Valko, Mohammad Ghavamzadeh, and Alessandro Lazaric · 2012
Cited alongside, same era.
A weakly supervised approach for object detection based on soft-label boosting
Weihong Wang, Yang Wang, Fang Chen, and Arcot Sowmya · 2013
Cited alongside, same era.
Positive-unlabeled learning with non-negative risk estimator
Ryuichi Kiryo, Gang Niu, Marthinus C du Plessis, and Masashi Sugiyama · 2017
Later among the works it cites.
InfoGAIL: Interpretable imitation learning from visual demonstrations
Yunzhu Li, Jiaming Song, and Stefano Ermon · 2017
Later among the works it cites.
Learning features by watching objects move
Deepak Pathak, Ross B Girshick, Piotr Dollár, Trevor Darrell, and Bharath Hariharan · 2017
Later among the works it cites.
A deep reinforcement learning chatbot
Iulian V Serban, Chinnadhurai Sankar, Mathieu Germain, Saizheng Zhang, Zhouhan Lin, Sandeep Subramanian, Taesup Kim, Michael Pieper, Sarath Chandar, Nan Rosemary Ke, et al · 2017
Later among the works it cites.
Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning
James Steven Supancic III and Deva Ramanan · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning and the reward engineering principle
Daniel Dewey · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Soft label based semi-supervised boosting for classification and object recognition
Dingfu Zhou, Benjamin Quost, and Vincent Frémont · 2014
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Distance minimization for reward learning from scored trajectories
Benjamin Burchfiel, Carlo Tomasi, and Ronald Parr · 2016
Cited alongside, same era.
Learning motion patterns in videos
Pavel Tokmakov, Karteek Alahari, and Cordelia Schmid · 2017
Later among the works it cites.
Learning to learn from noisy web videos
Serena Yeung, Vignesh Ramanathan, Olga Russakovsky, Liyue Shen, Greg Mori, and Li Fei-Fei · 2017
Later among the works it cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Later among the works it cites.
Inference aided reinforcement learning for incentive mechanism design in crowdsourcing
Zehong Hu, Yitao Liang, Jie Zhang, Zhao Li, and Yang Liu · 2018
Later among the works it cites.
Binary classification from positive-confidence data
Takashi Ishida, Gang Niu, and Masashi Sugiyama · 2018
Later among the works it cites.
Policy optimization with demonstrations
Bingyi Kang, Zequn Jie, and Jiashi Feng · 2018
Later among the works it cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning
Pararth Shah, Dilek Hakkani-Tur, Bing Liu, and Gokhan Tur · 2018
Later among the works it cites.
Addressing sample inefficiency and reward bias in inverse reinforcement learning
Ilya Kostrikov, Kumar Krishna Agrawal, Sergey Levine, and Jonathan Tompson · 2019
Closest in time.