Fetching the paper…
Reading the bibliography…
Behavior cloning (BC) is often practical for robot learning because it allows a policy to be trained offline without rewards, by supervised learning on expert demonstrations.
Reinforced imitation in heterogeneous action space
Konrad Zolna, Negar Rostamzadeh, Yoshua Bengio, Sungjin Ahn, and Pedro O Pinheiro · 1904
Earlier work this paper cites.
ALVINN: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2007
Earlier work this paper cites.
Learning classifiers from only positive and unlabeled data
Charles Elkan and Keith Noto · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Convex formulation for learning from positive and unlabeled data
Marthinus Du Plessis, Gang Niu, and Masashi Sugiyama · 2015
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
End-to-end differentiable adversarial imitation learning
Nir Baram, Oron Anschel, Itai Caspi, and Shie Mannor · 2017
Earlier work this paper cites.
Positive-unlabeled learning with non-negative risk estimator
Ryuichi Kiryo, Gang Niu, Marthinus C Du Plessis, and Masashi Sugiyama · 2017
Earlier work this paper cites.
InfoGAIL: Interpretable imitation learning from visual demonstrations
Yunzhu Li, Jiaming Song, and Stefano Ermon · 2017
Cited alongside, same era.
Learning human behaviors from motion capture by adversarial imitation
Josh Merel, Yuval Tassa, Sriram Srinivasan, Jay Lemmon, Ziyu Wang, Greg Wayne, and Nicolas Heess · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Večerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Cited alongside, same era.
Addressing sample inefficiency and reward bias in inverse reinforcement learning
BAIL: Best-action imitation learning for batch deep reinforcement learning
Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, Qing Deng, and Keith Ross · 2019
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
A practical approach to insertion with variable socket position using deep reinforcement learning
Mel Vecerik, Oleg Sushkov, David Barker, Thomas Rothörl, Todd Hester, and Jon Scholz · 2019
Later among the works it cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerik, Oleg Sushkov, David Barker, Jonathan Scholz, Misha Denil, Nando de Freitas, and Ziyu Wang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Kostrikov, Kumar Krishna Agrawal, Sergey Levine, and Jonathan Tompson · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Andrew Bagnell, Pieter Abbeel, and Jan Peters · 2018
Cited alongside, same era.
Observe and look further: Achieving consistent performance on Atari
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Večerík, et al · 2018
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Exponentially weighted imitation learning for batched historical data
Qing Wang, Jiechao Xiong, Lei Han, peng sun, Han Liu, and Tong Zhang · 2018
Cited alongside, same era.
Reinforcement and imitation learning for diverse visuomotor skills
Yuke Zhu, Ziyu Wang, Josh Merel, Andrei A. Rusu, Tom Erez, Serkan Cabi, Saran Tunyasuvunakool, János Kramár, Raia Hadsell, Nando de Freitas, and Nicolas Heess · 2018
Cited alongside, same era.
Closest in time.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Closest in time.
RL unplugged: Benchmarks for offline reinforcement learning
Çaglar Gülçehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gómez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel J. Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, Jerry Li, Mohammad Norouzi, Matthew D. Hoffman, Ofir Nachum, George Tucker, Nicolas Heess, and Nando de Freitas · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
SQIL: imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D. Dragan, and Sergey Levine · 2020
Closest in time.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
Noah Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Closest in time.
Critic regularized regression
Ziyu Wang, Alexander Novikov, Konrad Zolna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, et al · 2020
Closest in time.
Positive-unlabeled reward learning
Danfei Xu and Misha Denil · 2020
Closest in time.