Fetching the paper…
Reading the bibliography…
In recent years deep reinforcement learning (RL) systems have attained superhuman performance in a number of challenging task domains.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
The formation of learning sets
Harry F Harlow · 1949
Earlier work this paper cites.
A theory of pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement
Robert A Rescorla, Allan R Wagner, et al · 1972
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John C Gittins · 1979
Earlier work this paper cites.
Fixed-weight networks can learn
NE Cotter and PR Conwell · 1990
Earlier work this paper cites.
Simple principles of metalearning
Jurgen Schmidhuber, Jieyu Zhao, and Marco Wiering · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Learning to learn: Introduction and overview
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Fixed-weight on-line learning
A Steven Younger, Peter R Conwell, and Neil E Cotter · 1999
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A Steven Younger, and Peter R Conwell · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Adaptive behavior with fixed weights in rnn: an overview
Danil V Prokhorov, Lee A Feldkamp, and Ivan Yu Tyukin · 2002
Earlier work this paper cites.
Meta-learning in reinforcement learning
Nicolas Schweighofer and Kenji Doya · 2003
Earlier work this paper cites.
Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control
Nathaniel D Daw, Yael Niv, and Peter Dayan · 2005
Earlier work this paper cites.
Neural mechanism for stochastic behaviour during a competitive game
Alireza Soltani, Daeyeol Lee, and Xiao-Jing Wang · 2006
Earlier work this paper cites.
Learning the value of information in an uncertain world
Timothy EJ Behrens, Mark W Woolrich, Mark E Walton, and Matthew FS Rushworth · 2007
Cited alongside, same era.
Efficient learning in cellular simultaneous recurrent neural networks-the case of maze navigation problem
Roman Ilin, Robert Kozma, and Paul J Werbos · 2007
Cited alongside, same era.
Midbrain dopamine neurons signal preference for advance information about upcoming rewards
Ethan S Bromberg-Martin and Okihide Hikosaka · 2009
Cited alongside, same era.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Cited alongside, same era.
A meta-learning method based on temporal difference error
Kunikazu Kobayashi, Hiroyuki Mizoue, Takashi Kuremoto, and Masanao Obayashi · 2009
Cited alongside, same era.
Mechanisms for stochastic decision making in the primate frontal cortex: Single-neuron recording and circuit modeling
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Closest in time.
Learning to learn for global optimization of black box functions
Yutian Chen, Matthew W Hoffman, Sergio Gomez, Misha Denil, Timothy P Lillicrap, and Nando de Freitas · 2016
Closest in time.
Rl2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Closest in time.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daeyeol Lee and Xiao-Jing Wang · 2009
Cited alongside, same era.
Model-based influences on humans’ choices and striatal prediction errors
Nathaniel D Daw, Samuel J Gershman, Ben Seymour, Peter Dayan, and Raymond J Dolan · 2011
Cited alongside, same era.
Robot cognitive control with a neurophysiologically inspired reinforcement learning model
Mehdi Khamassi, Stéphane Lallée, Pierre Enel, Emmanuel Procyk, and Peter F Dominey · 2011
Cited alongside, same era.
Medial prefrontal cortex and the adaptive regulation of reinforcement learning parameters
Mehdi Khamassi, Pierre Enel, Peter Ford Dominey, and Emmanuel Procyk · 2013
Cited alongside, same era.
Bounded regret for finite-armed structured bandits
Tor Lattimore and Rémi Munos · 2014
Cited alongside, same era.
Learning to optimize via information-directed sampling
Dan Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Cited alongside, same era.
Max Jaderberg, Volodymir Mnih, Wojciech Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
When does model-based control pay off?
Wouter Kool, Fiery A Cushman, and Samuel J Gershman · 2016
Closest in time.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2016
Closest in time.
Ke Li and Jitendra Malik · 2016
Closest in time.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andy Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, Dharshan Kumaran, and Raia Hadsell · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, et al · 2016
Closest in time.
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Closest in time.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Closest in time.
The predictron: End-to-end learning and planning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, and Thomas Degris · 2017
Closest in time.
Meta-reinforcement learning: a bridge between prefrontal and dopaminergic function
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Joel Leibo, Hubert Soyer, Dharshan Kumaran, and Matthew Botvinick · 2017
Closest in time.