Fetching the paper…
Reading the bibliography…
Despite massive empirical evaluations, one of the fundamental questions in imitation learning is still not fully settled: does AIL (adversarial imitation learning) provably generalize better than BC (behavioral cloning)? We study this open problem with tabular and episodic MDPs.
An algorithm for quadratic programming
M. Frank, P. Wolfe, et al · 1956
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
D. Pomerleau · 1991
Earlier work this paper cites.
Mixed equilibria and dynamical systems arising from fictitious play in perturbed games
M. Benaım and M. W. Hirsch · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
T. Weissman, E. Ordentlich, G. Seroussi, S. Verdu, and M. J. Weinberger · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
L. Li, T. J. Walsh, and M. L. Littman · 2006
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
U. Syed and R. E. Schapire · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Dynamic Programming and Optimal Control: Volume I
D. Bertsekas · 2012
Earlier work this paper cites.
Online learning and online convex optimization
S. Shalev-Shwartz · 2012
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
S. Ross and J. A. Bagnell · 2014
Earlier work this paper cites.
Minimax estimation of discrete distributions under ℓ 1 \ell_{1} loss
Y. Han, J. Jiao, and T. Weissman · 2015
Earlier work this paper cites.
On learning distributions from their samples
S. Kamath, A. Orlitsky, D. Pichapati, and A. T. Suresh · 2015
Earlier work this paper cites.
Nonlinear Programming
D. P. Bertsekas · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Imitation learning: A survey of learning methods
A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne · 2017
Cited alongside, same era.
Learning robust rewards with adverserial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Cited alongside, same era.
Imitation learning via off-policy distribution matching
I. Kostrikov, O. Nachum, and J. Tompson · 2020
Later among the works it cites.
On gradient descent ascent for nonconvex-concave minimax problems
T. Lin, C. Jin, and M. I. Jordan · 2020
Later among the works it cites.
Toward the fundamental limits of imitation learning
N. Rajaraman, L. F. Yang, J. Jiao, and K. Ramchandran · 2020
Later among the works it cites.
Error bounds of imitating policies and environments
T. Xu, Z. Li, and Y. Yu · 2020
Later among the works it cites.
Apprenticeship learning via frank-wolfe
T. Zahavy, A. Cohen, H. Kaplan, and Y. Mansour · 2020
Later among the works it cites.
Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate
Y. Zhang, Q. Cai, Z. Yang, and Z. Wang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Cited alongside, same era.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, and J. Peters · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
Cited alongside, same era.
On the global convergence of imitation learning: A case for linear quadratic regulator
Q. Cai, M. Hong, Y. Chen, and Z. Wang · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
J. Chen and N. Jiang · 2019
Cited alongside, same era.
A divergence minimization perspective on imitation learning methods
S. K. S. Ghasemipour, R. S. Zemel, and S. Gu · 2019
Cited alongside, same era.
Near-optimal reward-free exploration for linear mixture mdps with plug-in solver
X. Chen, J. Hu, L. F. Yang, and L. Wang · 2021
Closest in time.
Primal wasserstein imitation learning
R. Dadashi, L. Hussenot, M. Geist, and O. Pietquin · 2021
Closest in time.
Iq-learn: Inverse soft-q learning for imitation
D. Garg, S. Chakraborty, C. Cundy, J. Song, and S. Ermon · 2021
Closest in time.
Adaptive reward-free exploration
E. Kaufmann, P. Ménard, O. D. Domingues, A. Jonsson, E. Leurent, and M. Valko · 2021
Closest in time.
Provably efficient generative adversarial imitation learning for online and offline setting with linear function approximation
Z. Liu, Y. Zhang, Z. Fu, Z. Yang, and Z. Wang · 2021
Closest in time.
Fast active learning for pure exploration in reinforcement learning
P. Ménard, O. D. Domingues, A. Jonsson, E. Kaufmann, E. Leurent, and M. Valko · 2021
Closest in time.
Of moments and matching: A game-theoretic framework for closing the imitation gap
G. Swamy, S. Choudhury, J. A. Bagnell, and S. Wu · 2021
Closest in time.
Representation learning for online and offline rl in low-rank mdps
M. Uehara, X. Zhang, and W. Sun · 2021
Closest in time.
Error bounds of imitating policies and environments for reinforcement learning
T. Xu, Z. Li, and Y. Yu · 2021
Closest in time.
Reward-free model-based reinforcement learning with linear function approximation
W. Zhang, D. Zhou, and Q. Gu · 2021
Closest in time.
Rethinking valuedice: Does it really improve performance?
Z. Li, T. Xu, Y. Yu, and Z.-Q. Luo · 2022
Closest in time.
Online apprenticeship learning
L. Shani, T. Zahavy, and S. Mannor · 2022
Closest in time.