Fetching the paper…
Reading the bibliography…
We present a novel algorithm to train a deep Q-learning agent using natural-gradient techniques.
On information and sufficiency
S. Kullback and R. A. Leibler · 1951
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain, 1994
J.R. Shewchuk · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2001
Earlier work this paper cites.
Natural Actor-Critic , pages 280–291
Jan Peters, Sethu Vijayakumar, and Stefan Schaal · 2005
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Cited alongside, same era.
MINRES-QLP: A krylov subspace method for indefinite or singular symmetric systems
Sou-Cheng T. Choi, Christopher C. Paige, and Michael A. Saunders · 2011
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Hamid Reza Maei · 2011
Cited alongside, same era.
The natural gradient, Jan 2013
Nick Foti · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Cited alongside, same era.
Prioritized Experience Replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Deep q-learning with natural gradients, Dec 2016
Alex Barron, Todor Markov, and Zack Swafford · 2016
Later among the works it cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Requests for research: Initial commit, 2016
OpenAI · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
Natural temporal difference learning
William Dabney and Philip S Thomas · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, and Koray Kavukcuoglu · 2015
Cited alongside, same era.
Lasagne: First release, August 2015
Sander Dieleman, Jan Schlüter, Colin Raffel, Eben Olson, Søren Kaae Sønderby, Daniel Nouri, Daniel Maturana, Martin Thoma, Eric Battenberg, Jack Kelly, Jeffrey De Fauw, Michael Heilman, Diogo Moitinho de Almeida, Brian McFee, Hendrik Weideman, Gábor Takács, Peter de Rivaz, Jon Crall, Gregory Sanders, Kashif Rasul, Cong Liu, Geoffrey French, and Jonas Degrave · 2015
Cited alongside, same era.
Natural conjugate gradient in variational inference
Antti Honkela, Matti Tornio, Tapani Raiko, and Juha Karhunen · 2015
Cited alongside, same era.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Later among the works it cites.
Yandex · 2016
Later among the works it cites.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2017
Later among the works it cites.
Deep Q-learning from Demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, G. Dulac-Arnold, I. Osband, J. Agapiou, J. Z. Leibo, and A. Gruslys · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
Y. Wu, E. Mansimov, S. Liao, R. Grosse, and J. Ba · 2017
Later among the works it cites.
Requests for research, 2018
OpenAI · 2018
Closest in time.
Winner’s curse? on pace, progress, and empirical rigor
D. Sculley, Jasper Snoek, Alex Wiltschko, and Ali Rahimi · 2018
Closest in time.