Fetching the paper…
Reading the bibliography…
This paper presents a novel form of policy gradient for model-free reinforcement learning (RL) with improved exploration properties.
Learning regular sets form queries and counterexamples
Dana Angulin · 1987
Earlier work this paper cites.
Some modified matrix eigenvalue problems
Gene Golub · 1987
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Efficient exploration in reinforcement learning
Sebastian B Thrun · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Learning in embedded systems
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Inductive Logic Programming: Theory and Methods
N. Lavrac and S. Dzeroski · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Artificial intelligence: a modern approach , volume 2
Stuart Jonathan Russell, Peter Norvig, John F Canny, Jitendra M Malik, and Douglas D Edwards · 2003
Earlier work this paper cites.
Optimal artificial curiosity, creativity, music, and the fine arts
Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Learning and using relational theories
Charles Kemp, Noah Goodman, and Joshua Tenebaum · 2007
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Episodic reinforcement learning by logistic reward-weighted regression
Daan Wierstra, Tom Schaul, Jan Peters, and Juergen Schmidhuber · 2008
Earlier work this paper cites.
Adaptive ε \varepsilon -greedy exploration in reinforcement learning based on value differences
Michel Tokic · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Cited alongside, same era.
Dynamic policy programming
Mohammad Gheshlaghi Azar, Vicenç Gómez, and Hilbert J Kappen · 2012
Cited alongside, same era.
Machine Learning: A Probabilistic Perspective
Kevin P. Murphy · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Cited alongside, same era.
Monte Carlo theory, methods and examples
Art B. Owen · 2013
Cited alongside, same era.
A new softmax operator for reinforcement learning
Kavosh Asadi and Michael L Littman · 2016
Closest in time.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Closest in time.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Closest in time.
G-learning: Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Closest in time.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwinska, Sergio G. Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, Adria P. Badia, Karl M. Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, and Demis Hassabis · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jörg Bornschein and Yoshua Bengio · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Cited alongside, same era.
Wojciech Zaremba and Ilya Sutskever · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, et al · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C. Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Closest in time.
Neural GPUs learn algorithms
Lukasz Kaiser and Ilya Sutskever · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Neural programmer: Inducing latent programs with gradient descent
Arvind Neelakantan, Quoc V. Le, and Ilya Sutskever · 2016
Closest in time.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans · 2016
Closest in time.
Deep exploration via bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Closest in time.
Neural programmer-interpreters
Scott E. Reed and Nando de Freitas · 2016
Closest in time.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Closest in time.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, et al · 2016
Closest in time.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Closest in time.