Fetching the paper…
Reading the bibliography…
We study Imitation Learning (IL) from Observations alone (ILFO) in large-scale MDPs.
Learning to predict by the methods of temporal differences
Sutton, R · 1988
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Müller, A · 1997
Earlier work this paper cites.
Predicting time series with support vector machines
Müller, K., Smola, A., and Rätsch, G · 1997
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Givan, R., Dean, T., and Greig, M · 2003
Earlier work this paper cites.
Exploration in metric state spaces
Kakade, S., Kearns, M. J., and Langford, J · 2003
Earlier work this paper cites.
Online Convex Programming and Generalized Infinitesimal Gradient Ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Policy search by dynamic programming
Bagnell, J. A., Kakade, S. M., Schneider, J. G., and Ng, A. Y · 2004
Earlier work this paper cites.
Distance-based classification with lipschitz functions
Luxburg, U. v. and Bousquet, O · 2004
Earlier work this paper cites.
Error limiting reductions between classification tasks
Beygelzimer, A., Dani, V., Hayes, T., Langford, J., and Zadrozny, B · 2005
Earlier work this paper cites.
Error bounds for approximate value iteration
Munos, R · 2005
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Search-based structured prediction
Daumé III, H., Langford, J., and Marcu, D · 2009
Earlier work this paper cites.
Kernel choice and classifiability for rkhs embeddings of probability distributions
Fukumizu, K., Gretton, A., Lanckriet, G. R., Schölkopf, B., and Sriperumbudur, B. K · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, J. A · 2010
Earlier work this paper cites.
A reduction from apprenticeship learning to classification
Syed, U. and Schapire, R. E · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G. J., and Bagnell, J · 2011
Cited alongside, same era.
Online learning and online convex optimization
Shalev-Shwartz, S. et al · 2012
Cited alongside, same era.
On the empirical estimation of integral probability metrics
Sriperumbudur, B. K., Fukumizu, K., Gretton, A., Schölkopf, B., Lanckriet, G. R., et al · 2012
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, A., Hsu, D., Kale, S., Langford, J., Li, L., and Schapire, R · 2014
Cited alongside, same era.
Analysis of classification-based policy iteration algorithms
Lazaric, A., Ghavamzadeh, M., and Munos, R · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D. et al · 2016
Later among the works it cites.
Learning to filter with predictive state inference machines
Sun, W., Venkatraman, A., Boots, B., and Bagnell, J. A · 2016
Later among the works it cites.
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Later among the works it cites.
Generalization and equilibrium in generative adversarial nets (gans)
Arora, S., Ge, R., Liang, Y., Ma, T., and Zhang, Y · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Cited alongside, same era.
Improving multi-step prediction of learned time series models
Venkatraman, A., Hebert, M., and Bagnell, J. A · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Dulac-Arnold, G., et al · 2017
Later among the works it cites.
Combining self-supervised learning and imitation for vision-based rope manipulation
Nair, A., Chen, D., Agrawal, P., Isola, P., Abbeel, P., Malik, J., and Levine, S · 2017
Later among the works it cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A · 2017
Later among the works it cites.
Imitating latent policies from observation
Edwards, A. D., Sahni, H., Schroeker, Y., and Isbell, C. L · 2018
Later among the works it cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Liu, Y., Gupta, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Agile autonomous driving using end-to-end deep imitation learning
Pan, Y., Cheng, C.-A., Saigol, K., Lee, K., Yan, X., Theodorou, E., and Boots, B · 2018
Later among the works it cites.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M · 2018
Later among the works it cites.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N · 2019
Closest in time.