Fetching the paper…
Reading the bibliography…
Currently, deep reinforcement learning (RL) shows impressive results in complex gaming and robotic environments.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
A. R. Cassandra, L. P. Kaelbling, and M. L. Littman · 1994
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Boosted bellman residual minimization handling expert demonstrations
B. Piot, M. Geist, and O. Pietquin · 2014
Earlier work this paper cites.
Boosted bellman residual minimization handling expert demonstrations
M. Piot, B.; Geist and O. Pietquin · 2014
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
D. S. V. Mnih, K. Kavukcuoglu and et al · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
A. G. H. Van Hasselt and D. Silver · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Playing atari games with deep reinforcement learning and human checkpoint replay
I. Hosu and T. Rebedea · 2016
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
A. P. M. M. G. A. L. T. H. T. S. D. Mnih, V.; Badia and K. Kavukcuoglu · 2016
Cited alongside, same era.
Prioritized experience replay
J. A. I. Schaul, T.; Quan and D. Silver · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
M. H. H. v. H. M. L. Ziyu Wang, Tom Schaul and N. de Freitas · 2016
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
One-shot imitation learning
Y. Duan, M. Andrychowicz, B. Stadie, O. J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine · 2018
Later among the works it cites.
Policy optimization with demonstrations
B. Kang, Z. Jie, and J. Feng · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Reinforcement and Imitation Learning for Diverse Visuomotor Skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, and N. Heess · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Relational reinforcement learning with guided demonstrations
D. Martínez, G. Alenya, and C. Torras · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Cited alongside, same era.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
T. Shu, C. Xiong, and R. Socher · 2017
Cited alongside, same era.
Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Cited alongside, same era.
Reinforcement learning from imperfect demonstrations
Y. Gao, J. Lin, F. Yu, S. Levine, T. Darrell, et al · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
T. Hester, M. Vecerík, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, G. Dulac-Arnold, J. Agapiou, J. Z. Leibo, and A. Gruslys
Cited in the paper.
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
B. Eysenbach, R. R. Salakhutdinov, and S. Levine · 2019
Later among the works it cites.
The MineRL competition on sample efficient reinforcement learning using human priors
W. H. Guss, C. Codel, K. Hofmann, B. Houghton, N. Kuno, S. Milani, S. Mohanty, D. P. Liebana, R. Salakhutdinov, N. Topin, et al · 2019
Later among the works it cites.
Making efficient use of demonstrations to solve hard exploration problems
T. L. Paine, C. Gulcehre, B. Shahriari, M. Denil, M. Hoffman, H. Soyer, R. Tanburn, S. Kapturowski, N. Rabinowitz, D. Williams, et al · 2019
Later among the works it cites.
Hierarchical Deep Q-Network from Imperfect Demonstrations in Minecraft
A. Skrynnik, A. Staroverov, E. Aitygulov, K. Aksenov, V. Davydov, and A. I. Panov · 2019
Later among the works it cites.
Improving Sample Efficiency in Model-Free Reinforcement Learning from Images
D. Yarats, A. Zhang, I. Kostrikov, B. o. Amos, J. Pineau, and R. Fergus · 2019
Later among the works it cites.