Fetching the paper…
Reading the bibliography…
We consider the problem of imitation learning under misspecification: settings where the learner is fundamentally unable to replicate expert behavior everywhere.
Problem complexity and method efficiency in optimization
Arkadij Semenovič Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Policy search by dynamic programming
James Bagnell, Sham M Kakade, Jeff Schneider, and Andrew Ng · 2003
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Maximum margin planning
Nathan D Ratliff, J Andrew Bagnell, and Martin A Zinkevich · 2006
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Umar Syed and Robert E Schapire · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Learning to search: Functional gradient techniques for imitation learning
Nathan D Ratliff, David Silver, and J Andrew Bagnell · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
Follow-the-regularized-leader and mirror descent: Equivalence theorems and l1 regularization
Brendan McMahan · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
On the universality of online mirror descent
Nati Srebro, Karthik Sridharan, and Ambuj Tewari · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V Mnih · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J Andrew Bagnell · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Constantinos Daskalakis, Andrew Ilyas, Vasilis Syrgkanis, and Haoyang Zeng · 2017
Cited alongside, same era.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Online apprenticeship learning
Lior Shani, Tom Zahavy, and Shie Mannor · 2022
Later among the works it cites.
Hybrid rl: Using both offline and online data can make rl efficient
Yuda Song, Yifei Zhou, Ayush Sekhari, J Andrew Bagnell, Akshay Krishnamurthy, and Wen Sun · 2022
Later among the works it cites.
Locomujoco: A comprehensive imitation learning benchmark for locomotion
Firas Al-Hafez, Guoping Zhao, Jan Peters, and Davide Tateo · 2023
Later among the works it cites.
Learning to generate better than your llm
Jonathan D Chang, Kiante Brantley, Rajkumar Ramamurthy, Dipendra Misra, and Wen Sun · 2023
Later among the works it cites.
Inverse reinforcement learning without reinforcement learning
Gokul Swamy, David Wu, Sanjiban Choudhury, Drew Bagnell, and Steven Wu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
Deep bayesian reward learning from preferences
Daniel S Brown and Scott Niekum · 2019
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2020
Cited alongside, same era.
Mitigating covariate shift in imitation learning via offline data with partial coverage
Jonathan Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Cited alongside, same era.
Mimicplay: Long-horizon imitation learning by watching human play
Chen Wang, Linxi Fan, Jiankai Sun, Ruohan Zhang, Li Fei-Fei, Danfei Xu, Yuke Zhu, and Anima Anandkumar · 2023
Later among the works it cites.
Provably efficient adversarial imitation learning with unknown transitions
Tian Xu, Ziniu Li, Yang Yu, and Zhi-Quan Luo · 2023
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2024
Later among the works it cites.
Dataset reset policy optimization for rlhf
Jonathan D Chang, Wenhao Zhan, Owen Oertell, Kianté Brantley, Dipendra Misra, Jason D Lee, and Wen Sun · 2024
Later among the works it cites.
Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning
Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris Kitani, Changliu Liu, and Guanya Shi · 2024
Later among the works it cites.
Sprinql: Sub-optimal demonstrations driven offline imitation learning
Huy Hoang, Tien Mai, and Pradeep Varakantham · 2024
Later among the works it cites.
When is agnostic reinforcement learning statistically tractable?
Zeyu Jia, Gene Li, Alexander Rakhlin, Ayush Sekhari, and Nati Srebro · 2024
Later among the works it cites.
Rapid motor adaptation for robotic manipulator arms
Yichao Liang, Kevin Ellis, and João Henriques · 2024
Later among the works it cites.
Inverse reinforcement learning with sub-optimal experts
Riccardo Poiani, Gabriele Curti, Alberto Maria Metelli, and Marcello Restelli · 2024
Later among the works it cites.
Hybrid inverse reinforcement learning
Juntao Ren, Gokul Swamy, Zhiwei Steven Wu, J Andrew Bagnell, and Sanjiban Choudhury · 2024
Later among the works it cites.
Evil: Evolution strategies for generalisable imitation learning
Silvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh, and Jakob Nicolaus Foerster · 2024
Later among the works it cites.
Imitation learning in discounted linear mdps without exploration assumptions
Luca Viano, Stratis Skoulakis, and Volkan Cevher · 2024
Later among the works it cites.
Wococo: Learning whole-body humanoid control with sequential contacts
Chong Zhang, Wenli Xiao, Tairan He, and Guanya Shi · 2024
Later among the works it cites.
Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills
Tairan He, Jiawei Gao, Wenli Xiao, Yuanhang Zhang, Zi Wang, Jiashun Wang, Zhengyi Luo, Guanqi He, Nikhil Sobanbab, Chaoyi Pan, et al · 2025
Closest in time.