Fetching the paper…
Reading the bibliography…
Consider learning a policy purely on the basis of demonstrated behavior -- that is, with no access to reinforcement signals, no knowledge of transition dynamics, and no further interaction with the environment.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1999
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang · 2006
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang · 2006
Earlier work this paper cites.
Apprenticeship learning using irl and gradient methods
Gergely Neu and Csaba Szepesvári · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Imitation learning with a value-based prior
Umar Syed and Robert E Schapire · 2007
Earlier work this paper cites.
Apprenticeship learning using irl and gradient methods
Gergely Neu and Csaba Szepesvári · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Imitation learning with a value-based prior
Umar Syed and Robert E Schapire · 2007
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Umar Syed and Robert E Schapire · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Training restricted boltzmann machines using approximations to the likelihood gradient
Tijmen Tieleman · 2008
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Umar Syed and Robert E Schapire · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Training restricted boltzmann machines using approximations to the likelihood gradient
Tijmen Tieleman · 2008
Earlier work this paper cites.
A reduction from apprenticeship learning to classification
Umar Syed and Robert E Schapire · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
Learning from demonstration using mdp induced metrics
Francisco S Melo and Manuel Lopes · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
A reduction from apprenticeship learning to classification
Umar Syed and Robert E Schapire · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
Learning from demonstration using mdp induced metrics
Francisco S Melo and Manuel Lopes · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Apprenticeship learning about multiple intentions
Monica Babes, Vukosi Marivate, and Michael L Littman · 2011
Earlier work this paper cites.
Map inference for bayesian inverse reinforcement learning
Jaedeug Choi and Kee-Eung Kim · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Batch, off-policy and model-free apprenticeship learning
Edouard Klein, Matthieu Geist, and Olivier Pietquin · 2011
Earlier work this paper cites.
Model-free apprenticeship learning for transfer of human impedance behaviour
Takeshi Mori, Matthew Howard, and Sethu Vijayakumar · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Apprenticeship learning about multiple intentions
Monica Babes, Vukosi Marivate, and Michael L Littman · 2011
Earlier work this paper cites.
Map inference for bayesian inverse reinforcement learning
Jaedeug Choi and Kee-Eung Kim · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Batch, off-policy and model-free apprenticeship learning
Edouard Klein, Matthieu Geist, and Olivier Pietquin · 2011
Earlier work this paper cites.
Model-free apprenticeship learning for transfer of human impedance behaviour
Takeshi Mori, Matthew Howard, and Sethu Vijayakumar · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Markov decision processes: methods and applications
Eugene A Feinberg and Adam Shwartz · 2012
Earlier work this paper cites.
Irl through structured classification
Edouard Klein, Matthieu Geist, Bilal Piot, and Olivier Pietquin · 2012
Earlier work this paper cites.
Markov decision processes: methods and applications
Eugene A Feinberg and Adam Shwartz · 2012
Earlier work this paper cites.
Irl through structured classification
Edouard Klein, Matthieu Geist, Bilal Piot, and Olivier Pietquin · 2012
Earlier work this paper cites.
Probabilistic inverse reinforcement learning in unknown environments
Aristide CY Tossou and Christos Dimitrakakis · 2013
Earlier work this paper cites.
Inverse reinforcement learning for compliant manipulation in letter handwriting
Ajay Kumar Tanwani and Aude Billard · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Probabilistic inverse reinforcement learning in unknown environments
Aristide CY Tossou and Christos Dimitrakakis · 2013
Earlier work this paper cites.
Inverse reinforcement learning for compliant manipulation in letter handwriting
Ajay Kumar Tanwani and Aude Billard · 2013
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Boosted and reward-regularized classification for apprenticeship learning
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Boosted and reward-regularized classification for apprenticeship learning
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
A bayesian approach to generative adversarial imitation learning
Wonseok Jeon, Seokin Seo, and Kee-Eung Kim · 2018
Later among the works it cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Later among the works it cites.
Imitation learning via kernel mean embedding
Kee-Eung Kim and Hyun Soo Park · 2018
Later among the works it cites.
Global overview of imitation learning
Alexandre Attia and Sharone Dayan · 2018
Later among the works it cites.
Rl baselines zoo
Antonin Raffin · 2018
Later among the works it cites.
Stable baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rlpy: a value-function-based reinforcement learning framework for education and research
Alborz Geramifard, Christoph Dann, Robert H Klein, William Dabney, and Jonathan P How · 2015
Cited alongside, same era.
Rlpy: a value-function-based reinforcement learning framework for education and research
Alborz Geramifard, Christoph Dann, Robert H Klein, William Dabney, and Jonathan P How · 2015
Cited alongside, same era.
Smooth imitation learning for online sequence prediction
Hoang M Le, Andrew Kang, Yisong Yue, and Peter Carr · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine · 2016
Cited alongside, same era.
Inverse reinforcement learning with simultaneous estimation of rewards and dynamics
Michael Herman, Tobias Gindele, Jörg Wagner, Felix Schmitt, and Wolfram Burgard · 2016
Cited alongside, same era.
Adversarial imitation via variational inverse reinforcement learning
Ahmed H Qureshi, Byron Boots, and Michael C Yip · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2019
Later among the works it cites.
Wasserstein adversarial imitation learning
Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche, Jalal Etesami, and Thai Hong Linh · 2019
Later among the works it cites.
Model-free irl using maximum likelihood estimation
Vinamra Jain, Prashant Doshi, and Bikramjit Banerjee · 2019
Later among the works it cites.
Truly batch apprenticeship learning with deep successor features
Donghun Lee, Srivatsan Srinivasan, and Finale Doshi-Velez · 2019
Later among the works it cites.
Sample-efficient imitation learning via gans
Lionel Blondé and Alexandros Kalousis · 2019
Later among the works it cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2019
Later among the works it cites.
Implicit generation and generalization in energy-based models
Yilun Du and Igor Mordatch · 2019
Later among the works it cites.
Openai gym: Rocket trajectory optimization is a classic topic in optimal control
Oleg Klimov · 2019
Later among the works it cites.
Batch apprenticeship learning
Donghun Lee, Srivatsan Srinivasan, and Finale Doshi-Velez · 2019
Later among the works it cites.
Random expert distillation: Imitation learning via expert policy support estimation
Ruohan Wang, Carlo Ciliberto, Pierluigi Amadori, and Yiannis Demiris · 2019
Later among the works it cites.
Adversarial imitation via variational inverse reinforcement learning
Ahmed H Qureshi, Byron Boots, and Michael C Yip · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2019
Later among the works it cites.
Wasserstein adversarial imitation learning
Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche, Jalal Etesami, and Thai Hong Linh · 2019
Later among the works it cites.
Model-free irl using maximum likelihood estimation
Vinamra Jain, Prashant Doshi, and Bikramjit Banerjee · 2019
Later among the works it cites.
Truly batch apprenticeship learning with deep successor features
Donghun Lee, Srivatsan Srinivasan, and Finale Doshi-Velez · 2019
Later among the works it cites.
Sample-efficient imitation learning via gans
Lionel Blondé and Alexandros Kalousis · 2019
Later among the works it cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2019
Later among the works it cites.
Implicit generation and generalization in energy-based models
Yilun Du and Igor Mordatch · 2019
Later among the works it cites.
Openai gym: Rocket trajectory optimization is a classic topic in optimal control
Oleg Klimov · 2019
Later among the works it cites.
Batch apprenticeship learning
Donghun Lee, Srivatsan Srinivasan, and Finale Doshi-Velez · 2019
Later among the works it cites.
Random expert distillation: Imitation learning via expert policy support estimation
Ruohan Wang, Carlo Ciliberto, Pierluigi Amadori, and Yiannis Demiris · 2019
Later among the works it cites.
Inverse active sensing: Modeling and understanding timely decision-making
Daniel Jarrett and Mihaela van der Schaar · 2020
Closest in time.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2020
Closest in time.
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2020
Closest in time.
On the anatomy of mcmc-based maximum likelihood learning of energy-based models
Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu · 2020
Closest in time.
Sqil: Imitation learning via regularized behavioral cloning
Siddharth Reddy, Anca D Dragan, and Sergey Levine · 2020
Closest in time.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2020
Closest in time.
Jem - joint energy models
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2020
Closest in time.
How to train your energy-based model for regression
Fredrik K Gustafsson, Martin Danelljan, Radu Timofte, and Thomas B Schön · 2020
Closest in time.
State alignment-based imitation learning
Fangchen Liu, Zhan Ling, Tongzhou Mu, and Hao Su · 2020
Closest in time.
Disagreement-regularized imitation learning
Kianté Brantley, Wen Sun, and Mikael Henaff · 2020
Closest in time.
Primal wasserstein imitation learning
Robert Dadashi, Leonard Hussenot, Matthieu Geist, and Olivier Pietquin · 2020
Closest in time.
Energy-based imitation learning
Minghuan Liu, Tairan He, Minkai Xu, and Weinan Zhang · 2020
Closest in time.
Inverse active sensing: Modeling and understanding timely decision-making
Daniel Jarrett and Mihaela van der Schaar · 2020
Closest in time.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2020
Closest in time.
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2020
Closest in time.
On the anatomy of mcmc-based maximum likelihood learning of energy-based models
Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu · 2020
Closest in time.
Sqil: Imitation learning via regularized behavioral cloning
Siddharth Reddy, Anca D Dragan, and Sergey Levine · 2020
Closest in time.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2020
Closest in time.
Jem - joint energy models
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2020
Closest in time.
How to train your energy-based model for regression
Fredrik K Gustafsson, Martin Danelljan, Radu Timofte, and Thomas B Schön · 2020
Closest in time.
State alignment-based imitation learning
Fangchen Liu, Zhan Ling, Tongzhou Mu, and Hao Su · 2020
Closest in time.
Disagreement-regularized imitation learning
Kianté Brantley, Wen Sun, and Mikael Henaff · 2020
Closest in time.
Primal wasserstein imitation learning
Robert Dadashi, Leonard Hussenot, Matthieu Geist, and Olivier Pietquin · 2020
Closest in time.
Energy-based imitation learning
Minghuan Liu, Tairan He, Minkai Xu, and Weinan Zhang · 2020
Closest in time.