Fetching the paper…
Reading the bibliography…
We tackle a common scenario in imitation learning (IL), where agents try to recover the optimal policy from expert demonstrations without further access to the expert or environment reward signals.
Implicit Generation and Generalization in Energy-Based Models
Yilun Du and Igor Mordatch. 2019a · 1903
Earlier work this paper cites.
Learning Non-Convergent Non-Persistent Short-Run MCMC Toward Energy-Based Model
Erik Nijkamp, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu. 2019 · 1904
Earlier work this paper cites.
SQIL: Imitation Learning via Regularized Behavioral Cloning
Siddharth Reddy, Anca D. Dragan, and Sergey Levine. 2019 · 1905
Earlier work this paper cites.
Random Expert Distillation: Imitation Learning via Expert Policy Support Estimation
Ruohan Wang, Carlo Ciliberto, Pierluigi Vito Amadori, and Yiannis Demiris. 2019 · 1905
Earlier work this paper cites.
Efficient Training of Artificial Neural Networks for Autonomous Navigation
Dean A Pomerleau. 1991 · 1991
Earlier work this paper cites.
Algorithms for inverse reinforcement learning.. In Icml , Vol. 1. 2
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning . 1
Pieter Abbeel and Andrew Y Ng. 2004 · 2004
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Brian Sallans and Geoffrey E Hinton. 2004 · 2004
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang. 2006 · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning.. In Aaai , Vol. 8. Chicago, IL, USA, 1433–1438
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey. 2008 · 2008
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics . 297–304
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning. In Proceedings of the thirteenth international conference on artificial intelligence and statistics . 661–668
Stéphane Ross and Drew Bagnell. 2010 · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart. 2010 · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics . 627–635
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent. 2011 · 2011
Earlier work this paper cites.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Michael U Gutmann and Aapo Hyvärinen. 2012 · 2012
Earlier work this paper cites.
Actor-Critic Reinforcement Learning with Energy-Based Policies.. In EWRL . 43–58
Nicolas Heess, David Silver, and Yee Whye Teh. 2012 · 2012
Cited alongside, same era.
Infinite time horizon maximum causal entropy inverse reinforcement learning. In 53rd IEEE Conference on Decision and Control . IEEE, 4911–4916
Michael Bloem and Nicholas Bambos. 2014 · 2014
Cited alongside, same era.
Generative adversarial nets. In Advances in neural information processing systems . 2672–2680
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Trust region policy optimization. In International conference on machine learning . 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Openai baselines (2017)
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov. 2016 · 2016
Cited alongside, same era.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan. 2018 · 2018
Later among the works it cites.
Exploration by Random Network Distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. 2018 · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Imitation learning via kernel mean embedding. In Thirty-Second AAAI Conference on Artificial Intelligence
Kee-Eung Kim and Hyun Soo Park. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine. 2016a · 2016
Cited alongside, same era.
Generative adversarial imitation learning. In Advances in neural information processing systems . 4565–4573
Jonathan Ho and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
f-gan: Training generative neural samplers using variational divergence minimization. In Advances in neural information processing systems . 271–279
Sebastian Nowozin, Botond Cseke, and Ryota Tomioka. 2016 · 2016
Cited alongside, same era.
Energy-based generative adversarial network
Junbo Zhao, Michael Mathieu, and Yann LeCun. 2016 · 2016
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Improved training of wasserstein gans. In Advances in neural information processing systems . 5767–5777
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. 2017 · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 1352–1361
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson. 2018 · 2018
Later among the works it cites.
Deep energy estimator networks
Saeed Saremi, Arash Mehrjou, Bernhard Schölkopf, and Aapo Hyvärinen. 2018 · 2018
Later among the works it cites.
Regularizing Model-Based Planning with Energy-Based Models
Rinu Boney, Juho Kannala, and Alexander Ilin. 2019 · 2019
Later among the works it cites.
A Divergence Minimization Perspective on Imitation Learning Methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu. 2019 · 2019
Later among the works it cites.
Imitation Learning as f f -Divergence Minimization
Liyiming Ke, Matt Barnes, Wen Sun, Gilwoo Lee, Sanjiban Choudhury, and Siddhartha Srinivasa. 2019 · 2019
Later among the works it cites.
State Alignment-based Imitation Learning
Fangchen Liu, Zhan Ling, Tongzhou Mu, and Hao Su. 2019 · 2019
Later among the works it cites.
Saeed Saremi and Aapo Hyvarinen. 2019 · 2019
Later among the works it cites.
Sliced score matching: A scalable approach to density and score estimation
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2019 · 2019
Later among the works it cites.
Provably efficient imitation learning from observation alone
Wen Sun, Anirudh Vemula, Byron Boots, and J Andrew Bagnell. 2019 · 2019
Later among the works it cites.
Disagreement-Regularized Imitation Learning. In International Conference on Learning Representations
Kiante Brantley, Wen Sun, and Mikael Henaff. 2020 · 2020
Closest in time.
Sliced score matching: A scalable approach to density and score estimation. In Uncertainty in Artificial Intelligence . PMLR, 574–584
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2020 · 2020
Closest in time.