Fetching the paper…
Reading the bibliography…
Adversarial imitation learning has become a popular framework for imitation in continuous control.
“Dropout: a simple way to prevent neural networks from overfitting”
Nitish Srivastava et al · 1958
Earlier work this paper cites.
“Comparing biases for minimal network construction with back-propagation”
Stephen Hanson and Lorien Pratt · 1988
Earlier work this paper cites.
“Efficient training of artificial neural networks for autonomous navigation”
Dean Pomerleau · 1991
Earlier work this paper cites.
“Learning agents for uncertain environments”
Stuart Russell · 1998
Earlier work this paper cites.
“Is imitation learning the route to humanoid robots?”
Stefan Schaal · 1999
Earlier work this paper cites.
“Algorithms for inverse reinforcement learning.”
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
“Maximum entropy inverse reinforcement learning.”
Brian Ziebart, Andrew Maas, J Bagnell and Anind Dey · 2008
Earlier work this paper cites.
“A survey of robot learning from demonstration”
Brenna Argall, Sonia Chernova, Manuela Veloso and Brett Browning · 2009
Earlier work this paper cites.
“Modeling purposeful adaptive behavior with the principle of maximum causal entropy”, 2010
Brian Ziebart · 2010
Earlier work this paper cites.
“Generative adversarial networks”
Ian Goodfellow et al · 2014
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“High-dimensional continuous control using generalized advantage estimation”
John Schulman et al · 2015
Earlier work this paper cites.
“Trust region policy optimization”
John Schulman et al · 2015
Earlier work this paper cites.
Greg Brockman et al · 2016
Earlier work this paper cites.
“Generative adversarial imitation learning”
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
“Mastering the game of Go with deep neural networks and tree search”
David Silver et al · 2016
Earlier work this paper cites.
“A distributional perspective on reinforcement learning”
Marc Bellemare, Will Dabney and Rémi Munos · 2017
Earlier work this paper cites.
“Learning robust rewards with adversarial inverse reinforcement learning”
Justin Fu, Katie Luo and Sergey Levine · 2017
Earlier work this paper cites.
“Improved training of wasserstein gans”
Ishaan Gulrajani et al · 2017
Cited alongside, same era.
“Reproducibility of benchmarked deep reinforcement learning tasks for continuous control”
Riashat Islam, Peter Henderson, Maziar Gomrokchi and Doina Precup · 2017
Cited alongside, same era.
“Data-efficient deep reinforcement learning for dexterous manipulation”
Ivaylo Popov et al · 2017
Cited alongside, same era.
“Learning complex dexterous manipulation with deep reinforcement learning and demonstrations”
Aravind Rajeswaran et al · 2017
Cited alongside, same era.
“The mirage of action-dependent baselines in reinforcement learning”
George Tucker et al · 2018
Later among the works it cites.
“Dota 2 with large scale deep reinforcement learning”
Christopher Berner et al · 2019
Later among the works it cites.
“Benchmarking model-based reinforcement learning”
Eric Langlois et al · 2019
Later among the works it cites.
“Making efficient use of demonstrations to solve hard exploration problems”
Tom Paine et al · 2019
Later among the works it cites.
“Grandmaster level in StarCraft II using multi-agent reinforcement learning”
Oriol Vinyals et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman et al · 2017
Cited alongside, same era.
“mixup: Beyond empirical risk minimization”
Hongyi Zhang, Moustapha Cisse, Yann Dauphin and David Lopez-Paz · 2017
Cited alongside, same era.
“Distributed distributional deterministic policy gradients”
Gabriel Barth-Maron et al · 2018
Cited alongside, same era.
“JAX: composable transformations of Python+NumPy programs”, 2018
James Bradbury et al · 2018
Cited alongside, same era.
“Addressing function approximation error in actor-critic methods”
Scott Fujimoto, Herke Hoof and David Meger · 2018
Cited alongside, same era.
“Soft actor-critic algorithms and applications”
Tuomas Haarnoja et al · 2018
Cited alongside, same era.
“Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor”
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel and Sergey Levine · 2018
Cited alongside, same era.
“Deep reinforcement learning that matters”
Peter Henderson et al · 2018
Cited alongside, same era.
“Positive-unlabeled reward learning”
Danfei Xu and Misha Denil · 2019
Later among the works it cites.
“What matters in on-policy reinforcement learning? a large-scale empirical study”
Marcin Andrychowicz et al · 2020
Later among the works it cites.
“Learning dexterous in-hand manipulation”
OpenAI: Andrychowicz et al · 2020
Later among the works it cites.
“Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning”
Lionel Blondé, Pablo Strasser and Alexandros Kalousis · 2020
Later among the works it cites.
“D4rl: Datasets for deep data-driven reinforcement learning”
Justin Fu et al · 2020
Later among the works it cites.
“A divergence minimization perspective on imitation learning methods”
Seyed Ghasemipour, Richard Zemel and Shixiang Gu · 2020
Later among the works it cites.
“Flax: A neural network library and ecosystem for JAX”, 2020
Jonathan Heek et al · 2020
Later among the works it cites.
“Optax: composable gradient transformation and optimisation, in JAX!”, 2020
Matteo Hessel et al · 2020
Later among the works it cites.
“Acme: A research framework for distributed reinforcement learning”
Matt Hoffman et al · 2020
Later among the works it cites.
“Scaling laws for neural language models”
Jared Kaplan et al · 2020
Later among the works it cites.
“Designing Network Design Spaces”
Ilija Radosavovic et al · 2020
Later among the works it cites.
“Batch exploration with examples for scalable robotic reinforcement learning”
Annie Chen, HyunJi Nam, Suraj Nair and Chelsea Finn · 2021
Closest in time.
“Hyperparameter Selection for Imitation Learning”, 2021
Leonard Hussenot et al · 2021
Closest in time.