Fetching the paper…
Reading the bibliography…
We propose a novel approach to train a multi-modal policy from mixed demonstrations without their behavior labels.
Efficient training of artificial neural networks for autonomous navigation
D. A. Pomerleau · 1991
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Robot programming by demonstration
A. Billard, S. Calinon, R. Dillmann, and S. Schaal · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Apprenticeship learning about multiple intentions
M. Babes, V. Marivate, K. Subramanian, and M. L. Littman · 2011
Earlier work this paper cites.
Bayesian multitask inverse reinforcement learning
C. Dimitrakakis and C. A. Rothkopf · 2011
Earlier work this paper cites.
Nonlinear inverse reinforcement learning with gaussian processes
S. Levine, Z. Popovic, and V. Koltun · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
TORCS, The Open Racing Car Simulator, v1.3.6, 2014
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner · 2014
Cited alongside, same era.
Learning structured output representation using deep conditional generative models
K. Sohn, H. Lee, and X. Yan · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Openai gym, 2016
Hierarchical attention networks for document classification
Z. Yang, D. Yang, C. Dyer, X. He, A. Smola, and E. Hovy · 2016
Later among the works it cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Later among the works it cites.
One-shot imitation learning
Y. Duan, M. Andrychowicz, B. Stadie, O. J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Later among the works it cites.
Multi-modal imitation learning from unstructured demonstrations using generative adversarial nets
K. Hausman, Y. Chebotar, S. Schaal, G. Sukhatme, and J. J. Lim · 2017
Later among the works it cites.
Infogail: Interpretable imitation learning from visual demonstrations
Y. Li, J. Song, and S. Ermon · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Tutorial on variational autoencoders
C. Doersch · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
C. J. Maddison, A. Mnih, and Y. W. Teh · 2016
Cited alongside, same era.
An uncertain future: Forecasting from static images using variational autoencoders
J. Walker, C. Doersch, A. Gupta, and M. Hebert · 2016
Cited alongside, same era.
J. Morton and M. J. Kochenderfer · 2017
Later among the works it cites.
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Robust imitation of diverse behaviors
Z. Wang, J. S. Merel, S. E. Reed, N. de Freitas, G. Wayne, and N. Heess · 2017
Later among the works it cites.
Imitation learning from visual data with multiple intentions
A. Tamar, K. Rohanimanesh, Y. Chow, C. Vigorito, B. Goodrich, M. Kahane, and D. Pridmore · 2018
Later among the works it cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Later among the works it cites.