Fetching the paper…
Reading the bibliography…
Existing imitation learning (IL) methods such as inverse reinforcement learning (IRL) usually have a double-loop training process, alternating between learning a reward function and a policy and tend to suffer long training time and high variance.
Efficient training of artificial neural networks for autonomous navigation
D. A. Pomerleau · 1991
Earlier work this paper cites.
Optimum design of Chamfer distance transforms
M. A. Butt and P. Maragos · 1998
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Boosting Chamfer matching by learning chamfer distance normalization
T. Ma, X. Yang, and L. J. Latecki · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
A. Boularias, J. Kober, and J. Peters · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Cited alongside, same era.
Sample efficient imitation learning for continuous control
F. Sasaki, T. Yohira, and A. Kawaguchi · 2018
Cited alongside, same era.
Taichi: a language for high-performance computation on spatially sparse data structures
Y. Hu, T.-M. Li, L. Anderson, J. Ragan-Kelley, and F. Durand · 2019
Cited alongside, same era.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
I. Kostrikov, K. K. Agrawal, D. Dwibedi, S. Levine, and J. Tompson · 2019
Cited alongside, same era.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2020
Later among the works it cites.
Off-policy imitation learning from observations
Z. Zhu, K. Lin, B. Dai, and J. Zhou · 2020
Later among the works it cites.
Ab Initio Particle-based Object Manipulation
S. Chen, X. Ma, Y. Lu, and D. Hsu · 2021
Later among the works it cites.
Primal Wasserstein imitation learning
R. Dadashi, L. Hussenot, M. Geist, and O. Pietquin · 2021
Later among the works it cites.
Brax – a differentiable physics engine for large scale rigid body simulation
C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem · 2021
Later among the works it cites.
Mastering atari with discrete world models
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Clavera, V. Fu, and P. Abbeel · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX, 2020
J. Heek, A. Levskaya, A. Oliver, M. Ritter, B. Rondepierre, A. Steiner, and M. van Zee · 2020
Cited alongside, same era.
Contrastive variational model-based reinforcement learning for complex observations
X. Ma, S. Chen, D. Hsu, and W. S. Lee · 2020
Cited alongside, same era.
Scalable differentiable physics for learning and control
Y.-L. Qiao, J. Liang, V. Koltun, and M. C. Lin · 2020
Cited alongside, same era.
MOPO: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Cited alongside, same era.
Later among the works it cites.
Disect: A differentiable simulation engine for autonomous robotic cutting
E. Heiden, M. Macklin, Y. Narang, D. Fox, A. Garg, and F. Ramos · 2021
Later among the works it cites.
Plasticinelab: A soft-body manipulation benchmark with differentiable physics
Z. Huang, Y. Hu, T. Du, S. Zhou, H. Su, J. B. Tenenbaum, and C. Gan · 2021
Later among the works it cites.
OPIRL: Sample efficient off-policy inverse reinforcement learning via distribution matching
H. Hoshino, K. Ota, A. Kanezaki, and R. Yokota · 2022
Closest in time.
Diffskill: Skill abstraction from differentiable physics for deformable object manipulations with tools
X. Lin, Z. Huang, Y. Li, J. B. Tenenbaum, D. Held, and C. Gan · 2022
Closest in time.