Fetching the paper…
Reading the bibliography…
Imitation learning practitioners have often noted that conditioning policies on previous actions leads to a dramatic divergence between "held out" error and performance of the learner in situ.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D · 1989
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Shimodaira, H · 2000
Earlier work this paper cites.
Causation, prediction, and search
Spirtes, P., Glymour, C. N., Scheines, R., and Heckerman, D · 2000
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Off-road obstacle avoidance through end-to-end learning
Muller, U., Ben, J., Cosatto, E., Flepp, B., and Cun, Y. L · 2006
Earlier work this paper cites.
Dynamic weighted majority: An ensemble method for drifting concepts
Kolter, J. Z. and Maloof, M. A · 2007
Earlier work this paper cites.
Machine learning techniques—reductions between prediction quality metrics
Beygelzimer, A., Langford, J., and Zadrozny, B · 2008
Earlier work this paper cites.
Mind the duality gap: Logarithmic regret algorithms for online optimization
Shalev-Shwartz, S. and Kakade, S. M · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, J. A · 2011
Cited alongside, same era.
Interactive learning for sequential decisions and predictions
Ross, S · 2013
Cited alongside, same era.
Machine learning: The high interest credit card of technical debt
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., and Young, M · 2014
Cited alongside, same era.
Feedback in machine learning
Bagnell, D · 2016
Cited alongside, same era.
End to end learning for self-driving cars
Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L. D., Monfort, M., Muller, U., Zhang, J., et al · 2016
Comparing human-centric and robot-centric sampling for robot deep learning from demonstrations
Laskey, M., Chuck, C., Lee, J., Mahler, J., Krishnan, S., Jamieson, K., Dragan, A., and Goldberg, K · 2017
Later among the works it cites.
Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst
Bansal, M., Krizhevsky, A., and Ogale, A · 2018
Later among the works it cites.
Disagreement-regularized imitation learning
Brantley, K., Sun, W., and Henaff, M · 2019
Later among the works it cites.
Exploring the limitations of behavior cloning for autonomous driving
Codevilla, F., Santana, E., López, A. M., and Gaidon, A · 2019
Later among the works it cites.
Causal confusion in imitation learning
de Haan, P., Jayaraman, D., and Levine, S · 2019
Later among the works it cites.
Provably efficient imitation learning from observation alone
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Causal inference in statistics: A primer
Pearl, J., Glymour, M., and Jewell, N. P · 2016
Cited alongside, same era.
Imitating driver behavior with generative adversarial networks
Kuefler, A., Morton, J., Wheeler, T., and Kochenderfer, M · 2017
Cited alongside, same era.
Sun, W., Vemula, A., Boots, B., and Bagnell, J. A · 2019
Later among the works it cites.
Adversarial soft advantage fitting: Imitation learning without policy optimization
Barde, P., Roy, J., Jeon, W., Pineau, J., Pal, C., and Nowrouzezahrai, D · 2020
Later among the works it cites.
Fighting copycat agents inbehavioral cloning from observation histories
Wen, C., Lin, J., Darrell, T., Jayaraman, D., and Gao, Y · 2020
Later among the works it cites.
Causal imitation learning with unobserved confounders
Zhang, J., Kumor, D., and Bareinboim, E · 2020
Later among the works it cites.