Fetching the paper…
Reading the bibliography…
Many modern methods for imitation learning and inverse reinforcement learning, such as GAIL or AIRL, are based on an adversarial formulation.
On information and sufficiency
S. Kullback and R. A. Leibler · 1951
Earlier work this paper cites.
Information Theory and Statistical Mechanics
E. T. Jaynes · 1957
Earlier work this paper cites.
Axiomatisations of the average and a further generalisation of monotonic sequences
J. Bibby · 1974
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time. iv
M. Donsker and S. Varadhan · 1983
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1989
Earlier work this paper cites.
Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Is Imitation Learning the Route to Humanoid Robots?
S. Schaal · 1999
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
A. Ng and S. Russell · 2000
Earlier work this paper cites.
Reducing the Time Complexity of the Derandomized Evolution Strategy with Covariance Matrix Adaptation (CMA-ES)
N. Hansen, S. Muller, and P. Koumoutsakos · 2003
Earlier work this paper cites.
An auxiliary variational method
F. V. Agakov and D. Barber · 2004
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
C. M. Bishop · 2006
Earlier work this paper cites.
Maximum Margin Planning
N. Ratliff, A. Bagnell, and M. Zinkevich · 2006
Earlier work this paper cites.
Direct importance estimation for covariate shift adaptation
M. Sugiyama, T. Suzuki, S. Nakajima, H. Kashima, P. von Bünau, and M. Kawanabe · 2008
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
B. Ziebart, A. Maas, A. Bagnell, and A. Dey · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Relative Entropy Inverse Reinforcement Learning
A. Boularias, J. Kober, and J. Peters · 2011
Earlier work this paper cites.
Nonlinear inverse reinforcement learning with gaussian processes
S. Levine, Z. Popovic, and V. Koltun · 2011
Cited alongside, same era.
Density-ratio matching under the bregman divergence: a unified framework of density-ratio estimation
M. Sugiyama, T. Suzuki, and T. Kanamori · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Cited alongside, same era.
Towards principled methods for training generative adversarial networks
M. Arjovsky and L. Bottou · 2017
Later among the works it cites.
Wasserstein generative adversarial networks
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Later among the works it cites.
Improved training of wasserstein gans
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville · 2017
Later among the works it cites.
Non-parametric policy search with limited information loss
H. van Hoof, G. Neumann, and J. Peters · 2017
Later among the works it cites.
Efficient gradient-free variational inference using policy search
O. Arenz, M. Zhong, and G. Neumann · 2018
Later among the works it cites.
Learning robust rewards with adverserial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Model-based relative entropy stochastic search
A. Abdolmaleki, R. Lioutikov, N. Lua, L. P. Reis, J. Peters, and G. Neumann · 2015
Cited alongside, same era.
Intent prediction and trajectory forecasting via predictive inverse linear-quadratic regulation
M. Monfort, A. Liu, and B. Ziebart · 2015
Cited alongside, same era.
Optimal control and inverse optimal control by distribution matching
O. Arenz, H. Abdulsamad, and G. Neumann · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Auxiliary deep generative models
L. Maaløe, C. K. Sønderby, S. K. Sønderby, and O. Winther · 2016
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
I. Kostrikov, K. K. Agrawal, D. Dwibedi, S. Levine, and J. Tompson · 2018
Later among the works it cites.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, J. Peters, et al · 2018
Later among the works it cites.
Generative adversarial imitation from observation
F. Torabi, G. Warnell, and P. Stone · 2018
Later among the works it cites.
Imitation learning as f f -divergence minimization
L. Ke, M. Barnes, W. Sun, G. Lee, S. Choudhury, and S. Srinivasa · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
O. Nachum, Y. Chow, B. Dai, and L. Li · 2019
Later among the works it cites.
Wasserstein adversarial imitation learning
H. Xiao, M. Herman, J. Wagner, S. Ziesche, J. Etesami, and T. H. Linh · 2019
Later among the works it cites.
Expected information maximization: Using the i-projection for mixture density estimation
P. Becker, O. Arenz, and G. Neumann · 2020
Closest in time.
A divergence minimization perspective on imitation learning methods
S. K. S. Ghasemipour, R. Zemel, and S. Gu · 2020
Closest in time.
Imitation learning via off-policy distribution matching
I. Kostrikov, O. Nachum, and J. Tompson · 2020
Closest in time.