Fetching the paper…
Reading the bibliography…
The inverse reinforcement learning approach to imitation learning is a double-edged sword.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1988
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
Policy search by dynamic programming
Bagnell, J. A., Kakade, S. M., Ng, A., and Schneider, J. G · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Ng, A. Y., Coates, A., Diel, M., Ganapathi, V., Schulte, J., Tse, B., Berger, E., and Liang, E · 2006
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E · 2007
Earlier work this paper cites.
A control architecture for quadruped locomotion over rough terrain
Kolter, J. Z., Rodgers, M. P., and Ng, A. Y · 2008
Earlier work this paper cites.
Learning to search: Functional gradient techniques for imitation learning
Ratliff, N. D., Silver, D., and Bagnell, J. A · 2009
Earlier work this paper cites.
Inverse optimal control with linearly-solvable mdps
Dvijotham, K. and Todorov, E · 2010
Earlier work this paper cites.
Learning from demonstration for autonomous navigation in complex unstructured terrain
Silver, D., Bagnell, J. A., and Stentz, A · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G. J., and Bagnell, J. A · 2011
Earlier work this paper cites.
Optimization and learning for rough terrain legged locomotion
Zucker, M., Ratliff, N., Stolle, M., Chestnutt, J., Bagnell, J. A., Atkeson, C. G., and Kuffner, J · 2011
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Ross, S. and Bagnell, J. A · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A · 2014
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Daskalakis, C., Ilyas, A., Syrgkanis, V., and Zeng, H · 2017
Cited alongside, same era.
Improved training of wasserstein gans
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C · 2017
Cited alongside, same era.
Learning robust rewards with adverserial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Instabilities of offline rl with pre-trained neural representation
Wang, R., Wu, Y., Salakhutdinov, R., and Kakade, S · 2021
Later among the works it cites.
Hierarchical model-based imitation learning for planning in autonomous driving
Bronstein, E., Palatucci, M., Notz, D., White, B., Kuefler, A., Lu, Y., Paul, S., Nikdel, P., Mougin, P., Chen, H., et al · 2022
Later among the works it cites.
Symphony: Learning realistic and diverse agents for autonomous driving simulation
Igl, M., Kim, D., Kuefler, A., Mougin, P., Shah, P., Shiarlis, K., Anguelov, D., Palatucci, M., White, B., and Whiteson, S · 2022
Later among the works it cites.
Hybrid rl: Using both offline and online data can make rl efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W · 2022
Later among the works it cites.
Minimax optimal online imitation learning via replay estimation
Swamy, G., Rajaraman, N., Peng, M., Choudhury, S., Bagnell, D., Wu, S., Jiao, J., and Ramchandran, K · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Cited alongside, same era.
Stable baselines3, 2019
Raffin, A., Hill, A., Ernestus, M., Gleave, A., Kanervisto, A., and Dormann, N · 2019
Cited alongside, same era.
Sqil: Imitation learning via regularized behavioral cloning
Reddy, S., Dragan, A. D., and Levine, S · 2019
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Xie, T. and Jiang, N · 2020
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
Du, S., Kakade, S., Lee, J., Lovett, S., Mahajan, G., Sun, W., and Wang, R · 2021
Cited alongside, same era.
Later among the works it cites.
Vinitsky, E., Lichtlé, N., Yang, X., Amos, B., and Foerster, J · 2022
Later among the works it cites.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Later among the works it cites.
Massively scalable inverse reinforcement learning in google maps, 2023
Barnes, M., Abueg, M., Lange, O. F., Deeds, M., Trader, J., Molitor, D., Wulfmeier, M., and O’Banion, S · 2023
Later among the works it cites.
Sequencematch: Imitation learning for autoregressive sequence modelling with backtracking
Cundy, C. and Ermon, S · 2023
Later among the works it cites.
Learning shared safety constraints from multi-task demonstrations
Kim, K., Swamy, G., Liu, Z., Zhao, D., Choudhury, S., and Wu, Z. S · 2023
Later among the works it cites.
Serl: A software suite for sample-efficient robotic reinforcement learning
Luo, J., Hu, Z., Xu, C., Gadipudi, S., Sharma, A., Ahmad, R., Schaal, S., Finn, C., Gupta, A., and Levine, S · 2023
Later among the works it cites.
Inverse reinforcement learning without reinforcement learning
Swamy, G., Choudhury, S., Bagnell, J. A., and Wu, Z. S · 2023
Later among the works it cites.
Tiapkin, D., Belomestny, D., Calandriello, D., Moulines, E., Naumov, A., Perrault, P., Valko, M., and Menard, P · 2023
Later among the works it cites.
The virtues of laziness in model-based rl: A unified objective and algorithms
Vemula, A., Song, Y., Singh, A., Bagnell, D., and Choudhury, S · 2023
Later among the works it cites.
Offline data enhanced on-policy policy gradient with provable guarantees
Zhou, Y., Sekhari, A., Song, Y., and Sun, W · 2023
Later among the works it cites.
Efficient imitation learning with conservative world models
Kolev, V., Rafailov, R., Hatch, K., Wu, J., and Finn, C · 2024
Closest in time.