Fetching the paper…
Reading the bibliography…
We propose a new framework for imitation learning -- treating imitation as a two-player ranking-based game between a policy and a reward.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
A general class of coefficients of divergence of one distribution from another
Syed Mumtaz Ali and Samuel D Silvey · 1966
Earlier work this paper cites.
Information-type measures of difference of probability distributions and indirect observation
Imre Csiszár · 1967
Earlier work this paper cites.
Note on discrimination information and variation (corresp.)
Igor Vajda · 1970
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1994
Earlier work this paper cites.
Dynamic noncooperative game theory
Tamer Başar and Geert Jan Olsder · 1998
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael Kearns and Satinder Singh · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
On the minimum f-divergence for given total variation
Gustavo Gilardoni · 2006
Earlier work this paper cites.
On divergences and informations in statistics and information theory
Friedrich Liese and Igor Vajda · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, and Roberto Martín-Martín · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
The submodular bregman and lovász-bregman divergences with applications: Extended version
Rishabh Iyer and Jeff Bilmes · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Reinforcement learning from demonstration through shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E Taylor, and Ann Nowé · 2015
Earlier work this paper cites.
Increasing the action gap: New operators for reinforcement learning
Marc G Bellemare, Georg Ostrovski, Arthur Guez, Philip Thomas, and Rémi Munos · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Cited alongside, same era.
Florian Schäfer and Anima Anandkumar · 2019
Later among the works it cites.
Provably efficient imitation learning from observation alone
Wen Sun, Anirudh Vemula, Byron Boots, and Drew Bagnell · 2019
Later among the works it cites.
Recent advances in imitation learning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2019
Later among the works it cites.
Imitation learning from observations by minimizing inverse dynamics disagreement
Chao Yang, Xiaojian Ma, Wenbing Huang, Fuchun Sun, Huaping Liu, Junzhou Huang, and Chuang Gan · 2019
Later among the works it cites.
Non-adversarial imitation learning and its connections to adversarial methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Least squares generative adversarial networks
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Oleg Arenz and Gerhard Neumann · 2020
Later among the works it cites.
Learning from suboptimal demonstration via self-supervised reward regression
Letian Chen, Rohan Paleja, and Matthew Gombolay · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Later among the works it cites.
Reinforcement learning from imperfect demonstrations under soft expert guidance
Mingxuan Jing, Xiaojian Ma, Wenbing Huang, Fuchun Sun, Chao Yang, Bin Fang, and Huaping Liu · 2020
Later among the works it cites.
When humans aren’t optimal: Robots that collaborate with risk-aware humans
Minae Kwon, Erdem Biyik, Aditi Talati, Karan Bhasin, Dylan P Losey, and Dorsa Sadigh · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
f-irl: Inverse reinforcement learning via state marginal matching
Tianwei Ni, Harshit Sikchi, Yufei Wang, Tejus Gupta, Lisa Lee, and Benjamin Eysenbach · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Later among the works it cites.
Self-guided actor-critic: Reinforcement learning from adaptive expert demonstrations
Haoran Zhang, Chenkun Yin, Yanxin Zhang, and Shangtai Jin · 2020
Later among the works it cites.
Learning reward functions from diverse sources of human feedback: Optimally integrating demonstrations and preferences
Erdem Bıyık, Dylan P Losey, Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh · 2021
Later among the works it cites.
Aux-airl: End-to-end self-supervised reward learning for extrapolating beyond suboptimal demonstrations
Yuchen Cui, Bo Liu, Akanksha Saran, Stephen Giguere, Peter Stone, and Scott Niekum · 2021
Later among the works it cites.
Iq-learn: Inverse soft-q learning for imitation
Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, and Stefano Ermon · 2021
Later among the works it cites.
Imitation learning as f-divergence minimization
Liyiming Ke, Sanjiban Choudhury, Matt Barnes, Wen Sun, Gilwoo Lee, and Siddhartha Srinivasa · 2021
Later among the works it cites.
Mobile: Model-based imitation learning from observation alone
Rahul Kidambi, Jonathan Chang, and Wen Sun · 2021
Later among the works it cites.
What matters for adversarial imitation learning?
Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent, Robert Dadashi, Sertan Girgin, Matthieu Geist, Olivier Bachem, Olivier Pietquin, and Marcin Andrychowicz · 2021
Later among the works it cites.
Of moments and matching: A game-theoretic framework for closing the imitation gap
Gokul Swamy, Sanjiban Choudhury, J Andrew Bagnell, and Steven Wu · 2021
Later among the works it cites.
Error bounds of imitating policies and environments for reinforcement learning
Tian Xu, Ziniu Li, and Yang Yu · 2021
Later among the works it cites.
Stackelberg actor-critic: Game-theoretic reinforcement learning algorithms
Liyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin Chasnov, and Lillian J Ratliff · 2021
Later among the works it cites.