Fetching the paper…
Reading the bibliography…
Often times in imitation learning (IL), the environment we collect expert demonstrations in and the environment we want to deploy our learned policy in aren't exactly the same (e.g.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1988
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Werbos, P · 1990
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Evolutionary algorithms in noisy environments: theoretical issues and guidelines for practice
Beyer, H.-G · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al · 2000
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A · 2001
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Ara*: Anytime a* with provable bounds on sub-optimality
Likhachev, M., Gordon, G. J., and Thrun, S · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Anytime dynamic a*: An anytime, replanning algorithm
Likhachev, M., Stentz, A., and Thrun, S · 2005
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Ng, A. Y., Coates, A., Diel, M., Ganapathi, V., Schulte, J., Tse, B., Berger, E., and Liang, E · 2006
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E · 2007
Earlier work this paper cites.
A control architecture for quadruped locomotion over rough terrain
Kolter, J. Z., Rodgers, M. P., and Ng, A. Y · 2008
Earlier work this paper cites.
Learning to search: Functional gradient techniques for imitation learning
Ratliff, N. D., Silver, D., and Bagnell, J. A · 2009
Earlier work this paper cites.
Genetic programming for reward function search
Niekum, S., Barto, A., and Spector, L · 2010
Earlier work this paper cites.
Artificial intelligence a modern approach
Russell, S. J. and Norvig, P · 2010
Earlier work this paper cites.
Learning from demonstration for autonomous navigation in complex unstructured terrain
Silver, D., Bagnell, J. A., and Stentz, A · 2010
Earlier work this paper cites.
Follow-the-regularized-leader and mirror descent: Equivalence theorems and l1 regularization
McMahan, B · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Optimization and learning for rough terrain legged locomotion
Zucker, M., Ratliff, N., Stolle, M., Chestnutt, J., Bagnell, J. A., Atkeson, C. G., and Kuffner, J · 2011
Cited alongside, same era.
Activity forecasting
Kitani, K. M., Ziebart, B. D., Bagnell, J. A., and Hebert, M · 2012
Cited alongside, same era.
Probabilistic pointing target prediction via inverse optimal control
Ziebart, B., Dey, A., and Bagnell, J. A · 2012
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models, 2016
Finn, C., Christiano, P., Abbeel, P., and Levine, S · 2016
Of moments and matching: A game-theoretic framework for closing the imitation gap
Swamy, G., Choudhury, S., Bagnell, J. A., and Wu, S · 2021
Later among the works it cites.
Fast population-based reinforcement learning on a single machine
Flajolet, A., Monroc, C. B., Beguir, K., and Pierrot, T · 2022
Later among the works it cites.
A theoretical understanding of gradient bias in meta-reinforcement learning, 2022
Liu, B., Feng, X., Ren, J., Mai, L., Zhu, R., Zhang, H., Wang, J., and Yang, Y · 2022
Later among the works it cites.
Gradients are not all you need, 2022
Metz, L., Freeman, C. D., Schoenholz, S. S., and Kachman, T · 2022
Later among the works it cites.
Minimax optimal online imitation learning via replay estimation
Swamy, G., Rajaraman, N., Peng, M., Choudhury, S., Bagnell, J., Wu, S. Z., Jiao, J., and Ramchandran, K · 2022
Later among the works it cites.
Nocturne: a scalable driving benchmark for bringing multi-agent learning one step closer to the real world
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generative adversarial imitation learning, 2016
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K · 2017
Cited alongside, same era.
Improved training of wasserstein gans, 2017
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning, 2017
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Training agent for first-person shooter game with actor-critic curriculum learning
Wu, Y. and Tian, Y · 2017
Cited alongside, same era.
Vinitsky, E., Lichtlé, N., Yang, X., Amos, B., and Foerster, J · 2022
Later among the works it cites.
Massively scalable inverse reinforcement learning in google maps
Barnes, M., Abueg, M., Lange, O. F., Deeds, M., Trader, J., Molitor, D., Wulfmeier, M., and O’Banion, S · 2023
Later among the works it cites.
Toward computationally efficient inverse reinforcement learning via reward shaping
Cooke, L. H., Klyne, H., Zhang, E., Laidlaw, C., Tambe, M., and Doshi-Velez, F · 2023
Later among the works it cites.
Discovering temporally-aware reinforcement learning algorithms
Jackson, M. T., Lu, C., Kirsch, L., Lange, R. T., Whiteson, S., and Foerster, J. N · 2023
Later among the works it cites.
Scaling opponent shaping to high dimensional games
Khan, A., Willi, T., Kwan, N., Tacchetti, A., Lu, C., Grefenstette, E., Rocktäschel, T., and Foerster, J · 2023
Later among the works it cites.
Bridging rl theory and practice with the effective horizon
Laidlaw, C., Russell, S., and Dragan, A · 2023
Later among the works it cites.
evosax: Jax-based evolution strategies
Lange, R. T · 2023
Later among the works it cites.
Adversarial cheap talk
Lu, C., Willi, T., Letcher, A., and Foerster, J. N · 2023
Later among the works it cites.
Hybrid inverse reinforcement learning
Ren, J., Swamy, G., Wu, Z. S., Bagnell, J. A., and Choudhury, S · 2023
Later among the works it cites.
Jaxmarl: Multi-agent rl environments in jax
Rutherford, A., Ellis, B., Gallici, M., Cook, J., Lupu, A., Ingvarsson, G., Willi, T., Khan, A., de Witt, C. S., Souly, A., et al · 2023
Later among the works it cites.
Inverse reinforcement learning without reinforcement learning, 2023
Swamy, G., Choudhury, S., Bagnell, J. A., and Wu, Z. S · 2023
Later among the works it cites.
Tiapkin, D., Belomestny, D., Calandriello, D., Moulines, E., Naumov, A., Perrault, P., Valko, M., and Menard, P · 2023
Later among the works it cites.
Waymax: An accelerated, data-driven simulator for large-scale autonomous driving research
Gulino, C., Fu, J., Luo, W., Tucker, G., Bronstein, E., Lu, Y., Harb, J., Pan, X., Wang, Y., Chen, X., et al · 2024
Closest in time.
Behaviour distillation
Lupu, A., Lu, C., Liesen, J. L., Lange, R. T., and Foerster, J. N · 2024
Closest in time.