Fetching the paper…
Reading the bibliography…
Imitation learning is a class of promising policy learning algorithms that is free from many practical issues with reinforcement learning, such as the reward design issue and the exploration hardness.
A framework for behavioural cloning
M. Bain and C. Sammut · 1995
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Visualizing data using t-sne
L. Van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas, A. Y. Hannun, A. Y. Ng, et al · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Dart: Noise injection for robust imitation learning
M. Laskey, J. Lee, R. Fox, A. Dragan, and K. Goldberg · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Earlier work this paper cites.
Policy optimization with demonstrations
B. Kang, Z. Jie, and J. Feng · 2018
Earlier work this paper cites.
Overcoming exploration in reinforcement learning with demonstrations
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Earlier work this paper cites.
Sample efficient imitation learning for continuous control
F. Sasaki, T. Yohira, and A. Kawaguchi · 2018
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
I. Kostrikov, K. K. Agrawal, D. Dwibedi, S. Levine, and J. Tompson · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Primal wasserstein imitation learning
R. Dadashi, L. Hussenot, M. Geist, and O. Pietquin · 2021
Later among the works it cites.
Mastering atari with discrete world models
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba · 2021
Later among the works it cites.
Learning and planning in complex action spaces
T. Hubert, J. Schrittwieser, I. Antonoglou, M. Barekatain, S. Schmitt, and D. Silver · 2021
Later among the works it cites.
Mobile: Model-based imitation learning from observation alone
R. Kidambi, J. Chang, and W. Sun · 2021
Later among the works it cites.
Model-based reinforcement learning via imagination with derived memory
Y. Mu, Y. Zhuang, B. Wang, G. Zhu, W. Liu, J. Chen, P. Luo, S. Li, C. Zhang, and J. Hao · 2021
Later among the works it cites.
What matters for adversarial imitation learning?
M. Orsini, A. Raichuk, L. Hussenot, D. Vincent, R. Dadashi, S. Girgin, M. Geist, O. Bachem, O. Pietquin, and M. Andrychowicz · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi · 2020
Cited alongside, same era.
Augmenting GAIL with BC for sample efficient imitation learning
R. Jena, C. Liu, and K. P. Sycara · 2020
Cited alongside, same era.
Imitation learning via off-policy distribution matching
I. Kostrikov, O. Nachum, and J. Tompson · 2020
Cited alongside, same era.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2020
Cited alongside, same era.
CURL: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Cited alongside, same era.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
A. X. Lee, A. Nagabandi, P. Abbeel, and S. Levine · 2020
Cited alongside, same era.
SQIL: imitation learning via reinforcement learning with sparse rewards
S. Reddy, A. D. Dragan, and S. Levine · 2020
Cited alongside, same era.
Later among the works it cites.
Learning long-term visual dynamics with region proposal interaction networks
H. Qi, X. Wang, D. Pathak, Y. Ma, and J. Malik · 2021
Later among the works it cites.
Visual adversarial imitation learning using variational models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Later among the works it cites.
Online and offline reinforcement learning by planning with a learned model
J. Schrittwieser, T. Hubert, A. Mandhane, M. Barekatain, I. Antonoglou, and D. Silver · 2021
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. C. Courville, and P. Bachman · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, R. D. Hjelm, P. Bachman, and A. C. Courville · 2021
Later among the works it cites.
Mastering atari games with limited data
W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao · 2021
Later among the works it cites.
Playvirtual: Augmenting cycle-consistent virtual trajectories for reinforcement learning
T. Yu, C. Lan, W. Zeng, M. Feng, Z. Zhang, and Z. Chen · 2021
Later among the works it cites.
Watch and match: Supercharging imitation with regularized optimal transport
S. Haldar, V. Mathur, D. Yarats, and L. Pinto · 2022
Closest in time.
Temporal difference learning for model predictive control
N. Hansen, X. Wang, and H. Su · 2022
Closest in time.
Rethinking valuedice: Does it really improve performance?
Z. Li, T. Xu, Y. Yu, and Z.-Q. Luo · 2022
Closest in time.
Mastering visual continuous control: Improved data-augmented reinforcement learning
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2022
Closest in time.
Spending thinking time wisely: Accelerating mcts with virtual expansions
W. Ye, P. Abbeel, and Y. Gao · 2022
Closest in time.