Fetching the paper…
Reading the bibliography…
Most existing imitation learning approaches assume the demonstrations are drawn from experts who are optimal, but relaxing this assumption enables us to use a wider range of data.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1995
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
Hiroaki Sugiyama, Toyomi Meguro, and Yasuhiro Minami · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Practical bilevel optimization: algorithms and applications
Jonathan F Bard · 2013
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Distance minimization for reward learning from scored trajectories
Benjamin Burchfiel, Carlo Tomasi, and Ronald Parr · 2016
Cited alongside, same era.
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Imitation learning from imperfect demonstration
Yueh-Hua Wu, Nontawat Charoenphakdee, Han Bao, Voot Tangkaratt, and Masashi Sugiyama · 2019
Later among the works it cites.
Erdem Bıyık, Dylan P Losey, Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh · 2020
Later among the works it cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Model-free preference-based reinforcement learning
Christian Wirth, Johannes Fürnkranz, and Gerhard Neumann · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Learning robust rewards with adverserial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning
Peter Henderson, Wei-Di Chang, Pierre-Luc Bacon, David Meger, Joelle Pineau, and Doina Precup · 2018
Cited alongside, same era.
Daniel S Brown, Wonjoon Goo, and Scott Niekum · 2020
Later among the works it cites.
Learning from suboptimal demonstration via self-supervised reward regression
Letian Chen, Rohan Paleja, and Matthew Gombolay · 2020
Later among the works it cites.
Dueling posterior sampling for preference-based reinforcement learning
Ellen R. Novoseller, Yibng Wei, Yanan Sui, Yisong Yue, and J. Burdick · 2020
Later among the works it cites.
Robust imitation learning from noisy demonstrations
Voot Tangkaratt, Nontawat Charoenphakdee, and Masashi Sugiyama · 2020
Later among the works it cites.
Variational imitation learning with diverse-quality demonstrations
Voot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, and Masashi Sugiyama · 2020
Later among the works it cites.
Causal imitation learning with unobserved confounders
Junzhe Zhang, Daniel Kumor, and Elias Bareinboim · 2020
Later among the works it cites.
Learning from imperfect demonstrations from agents with varying dynamics
Zhangjie Cao and Dorsa Sadigh · 2021
Closest in time.