Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al · 2018
Later among the works it cites.
Provably efficient imitation learning from observation alone
Wen Sun, Anirudh Vemula, Byron Boots, and Drew Bagnell · 2019
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Original
Siddharth Reddy, Anca D Dragan, and Sergey Levine · 2019
Later among the works it cites.
Stable baselines3, 2019
Antonin Raffin, Ashley Hill, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, and Noah Dormann · 2019
Later among the works it cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Original
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Of moments and matching: A game-theoretic framework for closing the imitation gap, 2021
Original
Gokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, and Zhiwei Steven Wu · 2021
Later among the works it cites.
Iq-learn: Inverse soft-q learning for imitation
Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, and Stefano Ermon · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
Hierarchical model-based imitation learning for planning in autonomous driving
Eli Bronstein, Mark Palatucci, Dominik Notz, Brandyn White, Alex Kuefler, Yiren Lu, Supratik Paul, Payam Nikdel, Paul Mougin, Hongge Chen, et al · 2022
Later among the works it cites.
Symphony: Learning realistic and diverse agents for autonomous driving simulation
Original
Maximilian Igl, Daewoo Kim, Alex Kuefler, Paul Mougin, Punit Shah, Kyriacos Shiarlis, Dragomir Anguelov, Mark Palatucci, Brandyn White, and Shimon Whiteson · 2022
Later among the works it cites.
Nocturne: a scalable driving benchmark for bringing multi-agent learning one step closer to the real world
Original
Eugene Vinitsky, Nathan Lichtlé, Xiaomeng Yang, Brandon Amos, and Jakob Foerster · 2022
Later among the works it cites.
Jump-start reinforcement learning
Original
Ikechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu, Mengyuan Yan, Joséphine Simon, Matthew Bennice, Chuyuan Fu, Cong Ma, Jiantao Jiao, et al · 2022
Later among the works it cites.
Minimax optimal online imitation learning via replay estimation
Original
Gokul Swamy, Nived Rajaraman, Matthew Peng, Sanjiban Choudhury, J Andrew Bagnell, Zhiwei Steven Wu, Jiantao Jiao, and Kannan Ramchandran · 2022
Later among the works it cites.
Modeling strong and human-like gameplay with kl-regularized search
Athul Paul Jacob, David J Wu, Gabriele Farina, Adam Lerer, Hengyuan Hu, Anton Bakhtin, Jacob Andreas, and Noam Brown · 2022
Later among the works it cites.
Regularized rl
Original
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines, Alexey Naumov, Pierre Perrault, Michal Valko, and Pierre Menard · 2023
Closest in time.