Fetching the paper…
Reading the bibliography…
We show that a critical vulnerability in adversarial imitation is the tendency of discriminator networks to learn spurious associations between visual features and expert labels.
Spurious correlation: A causal interpretation
H. A. Simon · 1954
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1989
Earlier work this paper cites.
Teaching by showing in kendama based on optimization principle, 1994
M. Kawato, F. Gandolfo, H. Gomi, and Y. Wada · 1994
Earlier work this paper cites.
Robot see, robot do: An overview of robot imitation
P. Bakker and Y. Kuniyoshi · 1996
Earlier work this paper cites.
A kendama learning robot based on bi-directional theory
H. Miyamoto, S. Schaal, F. Gandolfo, H. Gomi, Y. Koike, R. Osu, E. Nakano, Y. Wada, and M. Kawato · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. Bradtke and A. Barto · 1996
Earlier work this paper cites.
Learning from demonstration
S. Schaal · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
J. Pearl · 2009
Earlier work this paper cites.
An introduction to causal inference
J. Pearl · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
D. Ulyanov, A. Vedaldi, and V. Lempitsky · 2016
Cited alongside, same era.
Learning human behaviors from motion capture by adversarial imitation
J. Merel, Y. Tassa, S. Srinivasan, J. Lemmon, Z. Wang, G. Wayne, and N. Heess · 2017
Cited alongside, same era.
Elements of Causal Inference: Foundations and Learning Algorithms
J. Peters, D. Janzing, and B. Schölkopf · 2017
Cited alongside, same era.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, and N. Heess · 2018
Later among the works it cites.
X. B. Peng, A. Kanazawa, S. Toyer, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Visual imitation with a minimal adversary, 2018
S. Reed, Y. Aytar, Z. Wang, T. Paine, A. van den Oord, T. Pfaff, S. Gomez, A. Novikov, D. Budden, and O. Vinyals · 2018
Later among the works it cites.
Sample-efficient imitation learning via generative adversarial nets
L. Blondé and A. Kalousis · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One-shot imitation learning
Y. Duan, M. Andrychowicz, B. Stadie, O. J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. A. Riedmiller · 2017
Cited alongside, same era.
Causal effect inference with deep latent-variable models
C. Louizos, U. Shalit, J. M. Mooij, D. A. Sontag, R. S. Zemel, and M. Welling · 2017
Cited alongside, same era.
Infogail: Interpretable imitation learning from visual demonstrations
Y. Li, J. Song, and S. Ermon · 2017
Cited alongside, same era.
End-to-end differentiable adversarial imitation learning
N. Baram, O. Anschel, I. Caspi, and S. Mannor · 2017
Cited alongside, same era.
Third-person imitation learning
B. C. Stadie, P. Abbeel, and I. Sutskever · 2017
Cited alongside, same era.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, A. Muldal, N. Heess, and T. Lillicrap · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Later among the works it cites.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, A. Muldal, N. Heess, and T. Lillicrap · 2018
Later among the works it cites.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. Van Hasselt, and D. Silver · 2018
Later among the works it cites.
Large scale GAN training for high fidelity natural image synthesis
A. Brock, J. Donahue, and K. Simonyan · 2019
Closest in time.
Causal confusion in imitation learning
P. de Haan, D. Jayaraman, and S. Levine · 2019
Closest in time.
Off-policy evaluation in partially observable environments
G. Tennenholtz, S. Mannor, and U. Shalit · 2019
Closest in time.
Learning from demonstration in the wild
F. Behbahani, K. Shiarlis, X. Chen, V. Kurin, S. Kasewa, C. Stirbu, J. Gomes, S. Paul, F. A. Oliehoek, J. V. Messias, and S. Whiteson · 2019
Closest in time.
Sample efficient imitation learning for continuous control
F. Sasaki, T. Yohira, and A. Kawaguchi · 2019
Closest in time.
End-to-end robotic reinforcement learning without reward engineering
A. Singh, L. Yang, K. Hartikainen, C. Finn, and S. Levine · 2019
Closest in time.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
I. Kostrikov, K. K. Agrawal, D. Dwibedi, S. Levine, and J. Tompson · 2019
Closest in time.
Reinforced imitation in heterogeneous action space
K. Zolna, N. Rostamzadeh, Y. Bengio, S. Ahn, and P. O. Pinheiro · 2019
Closest in time.
Visual imitation learning with recurrent siamese networks
G. Berseth and C. J. Pal · 2019
Closest in time.