Fetching the paper…
Reading the bibliography…
We propose a new framework for Imitation Learning (IL) via density estimation of the expert's occupancy measure followed by Maximum Occupancy Entropy Reinforcement Learning (RL) using the density as a reward.
Iterative solution of games by fictitious play
George Brown · 1951
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
Estimation of Non-Normalized statistical models by score matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Sparse nonparametric density estimation in high dimensions using the rodeo
Han Liu, John Lafferty, and Larry Wasserman · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J. Wainwright, and Michael I. Jordan · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E Todorov, T Erez, and Y Tassa · 2012
Earlier work this paper cites.
Learning monocular reactive uav control in cluttered natural environments
Stephane Ross, Narek Melik-Barkhudarov, Kumar Shaurya Shankar, Andreas Wendel, Debadeepta Dey, Andrew J. Bagnell, and Martial Hebert · 2013
Earlier work this paper cites.
RNADE: The real-valued neural autoregressive density-estimator
Benigno Uria, Iain Murray, and Hugo Larochelle · 2013
Earlier work this paper cites.
NICE: Non-linear independent components estimation
L Dinh, D Krueger, and Y Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Xiaoxiao Guo, Satinder Singh, Honglak Lee, Richard L. Lewis, and Xiaoshi Wang · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in mujoco
Emanuel Todorov · 2014
Earlier work this paper cites.
Made: Masked autoencoder for distribution estimation
Mathieu Germain, Karol Gregor, Iain Murray, and Hugo Larochelle · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Zaremba Wojciech · 2016
Earlier work this paper cites.
Density estimation using real NVP
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Learning in implicit generative models
Shakir Mohamed and Balaji Lakshminarayanan · 2016
Cited alongside, same era.
f-GAN: Training generative neural samplers using variational divergence minimization
Sebastian Nowozin, Botond Cseke, and Ryota Tomioka · 2016
Cited alongside, same era.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Cited alongside, same era.
Neural autoregressive distribution estimation
Benigno Uria, Marc-Alexandre Côté, Karol Gregor, Iain Murray, and Hugo Larochelle · 2016
Cited alongside, same era.
Pixel recurrent neural networks
Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Later among the works it cites.
Multi-agent generative adversarial imitation learning
Jiaming Song, Hongyu Ren, Dorsa Sadigh, and Stefano Ermon · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron Van Den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Prescribed generative adversarial networks
Adji Dieng, Francisco Ruiz, David M. Blei, and Michalis K. Titsias · 2019
Later among the works it cites.
Implicit generation and generalization in energy-based models
Yilun Du and Igor Mordatch · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Variational inference using implicit distributions
Ferenc Huszár · 2017
Cited alongside, same era.
Masked autoregressive flow for density estimation
George Papamakarios, Theo Pavlakou, and Iain Murray · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D. Dragan, and Sergey Levine · 2017
Cited alongside, same era.
Robust imitation of diverse behaviors
Ziyu Wang, Josh Merel, Scott Reed, Greg Wayne, Nando de Freitas, and Nicolas Heess · 2017
Cited alongside, same era.
Later among the works it cites.
Flow contrastive estimation of energy-based models
Ruiqi Gao, Erik Nijkamp, Diederik P Kingma, Zhen Xu, Andrew M Dai, and Ying Nian Wu · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Entropy regularization with discounted future state distribution in policy gradient methods
Riashat Islam, Raihan Seraj, Pierre-Luc Bacon, and Doina Precup · 2019
Later among the works it cites.
What is local optimality in nonconvex-nonconcave minimax optimization?
Chi Jin, Praneeth Netrapalli, and Michael I. Jordan · 2019
Later among the works it cites.
Domain adaptive imitation learning
Kuno Kim, Yihong Gu, Jiaming Song, Shengjia Zhao, and Stefano Ermon · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Later among the works it cites.
Sliced score matching: A scalable approach to density and score estimation
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon · 2019
Later among the works it cites.
Random expert distillation: Imitation learning via expert policy support estimation
Ruohan Wang, Carlo Ciliberto, Pierluigi Amadori, and Yiannis Demiris · 2019
Later among the works it cites.
Model-based behavioral cloning with future image similarity learning
Alan Wu, AJ Piergiovanni, and Michael S. Ryoo · 2019
Later among the works it cites.
Disagreement-regularized imitation learning
Kiante Brantley, Wen Sun, and Mikael Henaff · 2020
Closest in time.
Imitation learning as f-divergence minimization
Liyiming Ke, Matt Barnes, Wen Sun, Gilwoo Lee, Sanjiban Choudhury, and Siddhartha Srinivasa · 2020
Closest in time.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2020
Closest in time.
Training deep energy-based models with f-divergence minimization
Lantao Yu, Yang Song, Jiaming Song, and Stefano Ermon · 2020
Closest in time.