Fetching the paper…
Reading the bibliography…
Recent progress in Reinforcement Learning (RL), fueled by its combination, with Deep Learning has enabled impressive results in learning to interact with complex virtual environments, yet real-world applications of RL are still scarce.
Switchboard: Telephone speech corpus for research and development
J. J. Godfrey, E. C. Holliman, and J. McDaniel · 1992
Earlier work this paper cites.
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Learning from demonstration
S. Schaal · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Frame skip is a powerful parameter for learning to play atari
A. Braylan, M. Hollenbeck, E. Meyerson, and R. Miikkulainen · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Chainer: a next-generation open source framework for deep learning
S. Tokui, K. Oono, S. Hido, and J. Clayton · 2015
Cited alongside, same era.
Model-based adversarial imitation learning
N. Baram, O. Anschel, and S. Mannor · 2016
Cited alongside, same era.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Playing atari games with deep reinforcement learning and human checkpoint replay
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
Inverse reinforcement learning from failure
K. Shiarlis, J. Messias, and S. Whiteson · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Later among the works it cites.
Exploration from demonstration for interactive reinforcement learning
K. Subramanian, C. L. Isbell Jr, and A. L. Thomaz · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I.-A. Hosu and T. Rebedea · 2016
Cited alongside, same era.
The malmo platform for artificial intelligence experimentation
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell · 2016
Cited alongside, same era.
Dynamic frame skip deep Q network
A. S. Lakshminarayanan, S. Sharma, and B. Ravindran · 2016
Cited alongside, same era.
A deep learning approach for joint video frame and reward prediction in atari games
F. Leibfried, N. Kushman, and K. Hofmann · 2016
Cited alongside, same era.
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, A. Sendonaris, G. Dulac-Arnold, I. Osband, J. Agapiou, J. Z. Leibo, and A. Gruslys · 2017
Closest in time.
Asynchronous data aggregation for training end to end visual control networks
M. Monfort, M. Johnson, A. Oliva, and K. Hofmann · 2017
Closest in time.
Learning to repeat: Fine grained action repetition for deep reinforcement learning
S. Sharma, A. S. Lakshminarayanan, and B. Ravindran · 2017
Closest in time.
Human learning in atari
P. A. Tsividis, T. Pouncy, J. L. Xu, J. B. Tenenbaum, and S. J. Gershman · 2017
Closest in time.