Fetching the paper…
Reading the bibliography…
Deep reinforcement learning is poised to revolutionise the field of AI and represents a step towards building autonomous systems with a higher level understanding of the visual world.
On the Theory of Dynamic Programming
Richard Bellman · 1952
Earlier work this paper cites.
Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences
Paul John Werbos · 1974
Earlier work this paper cites.
Neuronlike Adaptive Elements That Can Solve Difficult Learning Control Problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Asymptotically Efficient Adaptive Allocation Rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Learning Representations by Back-Propagating Errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1988
Earlier work this paper cites.
ALVINN, an Autonomous Land Vehicle in a Neural Network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Likelihood Ratio Gradient Estimation for Stochastic Systems
Peter W Glynn · 1990
Earlier work this paper cites.
Efficient Memory-Based Learning for Robot Control
Andrew William Moore · 1990
Earlier work this paper cites.
A Possibility for Implementing Curiosity and Boredom in Model-Building Neural Controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Learning to Generate Artificial Fovea Trajectories for Target Detection
Jürgen Schmidhuber and Rudolf Huber · 1991
Earlier work this paper cites.
Self-Improving Reactive Agents Based on Reinforcement Learning, Planning and Teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Q-Learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
Ronald J Williams · 1992
Earlier work this paper cites.
Advantage Updating
Leemon C Baird III · 1993
Earlier work this paper cites.
Improving Generalization for Temporal Difference Learning: The Successor Representation
Peter Dayan · 1993
Earlier work this paper cites.
On-line Q-learning using Connectionist Systems
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Temporal Difference Learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Multi-Player Residual Advantage Learning with General Function Approximation
Mance E Harmon and Leemon C Baird III · 1996
Earlier work this paper cites.
Multitask Learning
Rich Caruana · 1997
Earlier work this paper cites.
Analysis of Temporal-Difference Learning with Function Approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Planning and Acting in Partially Observable Stochastic Domains
Leslie P Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Deep Blue
Murray Campbell, A Joseph Hoane, and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
A Natural Policy Gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Optimizing Dialogue Management with Reinforcement Learning: Experiments with the NJFun System
Satinder Singh, Diane Litman, Michael Kearns, and Marilyn Walker · 2002
Earlier work this paper cites.
Covariant Policy Search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
On Actor-Critic Algorithms
Vijay R Konda and John N Tsitsiklis · 2003
Earlier work this paper cites.
Policy Gradient Reinforcement Learning for Fast Quadrupedal Locomotion
Nate Kohl and Peter Stone · 2004
Earlier work this paper cites.
Dynamic Programming and Suboptimal Control: A Survey from ADP to MPC
Dimitri P Bertsekas · 2005
Earlier work this paper cites.
Evolving Modular Fast-Weight Networks for Control
Faustino Gomez and Jürgen Schmidhuber · 2005
Earlier work this paper cites.
Path Integrals and Symmetry Breaking for Optimal Control Theory
Hilbert J Kappen · 2005
Earlier work this paper cites.
Neural Fitted Q Iteration—First Experiences with a Data Efficient Neural Reinforcement Learning Method
Martin Riedmiller · 2005
Earlier work this paper cites.
Gradient Estimation
Michael C Fu · 2006
Earlier work this paper cites.
Autonomous Inverted Helicopter Flight via Reinforcement Learning
Andrew Y Ng, Adam Coates, Mark Diel, Varun Ganapathi, Jamie Schulte, Ben Tse, Eric Berger, and Eric Liang · 2006
Earlier work this paper cites.
PAC Model-Free Reinforcement Learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
A Comprehensive survey of Multiagent Reinforcement Learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Managing Power Consumption and Performance of Computing Systems using Reinforcement Learning
Gerald Tesauro, Rajarshi Das, Hoi Chan, Jeffrey Kephart, David Levine, Freeman Rawson, and Charles Lefurgy · 2008
Earlier work this paper cites.
Curriculum Learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Building Watson: An Overview of the DeepQA Project
David Ferrucci, Eric Brown, Jennifer Chu-Carroll, James Fan, David Gondek, Aditya A Kalyanpur, Adam Lally, J William Murdock, Eric Nyberg, John Prager, et al · 2010
Earlier work this paper cites.
Relative Entropy Policy Search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Earlier work this paper cites.
Double Q-Learning
Hado van Hasselt · 2010
Earlier work this paper cites.
Recurrent Policy Gradients
Daan Wierstra, Alexander Förster, Jan Peters, and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Intrinsically Motivated Neuroevolution for Vision-Based Reinforcement Learning
Giuseppe Cuccu, Matthew Luciw, Jürgen Schmidhuber, and Faustino Gomez · 2011
Earlier work this paper cites.
Hogwild: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Stéphane Ross, Geoffrey J Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Large Scale Distributed Deep Networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Autonomous Reinforcement Learning on Raw Visual Input Data in a Real World Application
Sascha Lange, Martin Riedmiller, and Arne Voigtlander · 2012
Earlier work this paper cites.
MuJoCo: A Physics Engine for Model-Based Control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Representation Learning: A Review and New Perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Cited alongside, same era.
A Survey on Policy Search for Robotics
Marc P Deisenroth, Gerhard Neumann, and Jan Peters · 2013
Cited alongside, same era.
Evolving Large-Scale Neural Networks for Vision-Based Reinforcement Learning
Jan Koutník, Giuseppe Cuccu, Jürgen Schmidhuber, and Faustino Gomez · 2013
Cited alongside, same era.
Guided Policy Search
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2014
Cited alongside, same era.
Deep Exploration via Bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Later among the works it cites.
Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
Emilio Parisotto, Jimmy L Ba, and Ruslan Salakhutdinov · 2016
Later among the works it cites.
Sequence Level Training with Recurrent Neural Networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Later among the works it cites.
Prioritized Experience Replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Later among the works it cites.
High-Dimensional Continuous Control using Generalized Advantage Estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Later among the works it cites.
Taking the Human out of the Loop: A Review of Bayesian Optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando de Freitas · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Auto-Encoding Variational Bayes
Diederik P Kingma and Max Welling · 2014
Cited alongside, same era.
Learning Neural Network Policies with Guided Policy Search under Unknown Dynamics
Sergey Levine and Pieter Abbeel · 2014
Cited alongside, same era.
Recurrent Models of Visual Attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, and Koray Kavukcuoglu · 2014
Cited alongside, same era.
Stochastic Backpropagation and Approximate Inference in Deep Generative Models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Deterministic Policy Gradient Algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2015
Cited alongside, same era.
Later among the works it cites.
Mastering the Game of Go with Deep Neural Networks and Tree Search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Learning Multiagent Communication with Backpropagation
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2016
Later among the works it cites.
TorchCraft: A Library for Machine Learning Research on Real-Time Strategy Games
Gabriel Synnaeve, Nantas Nardelli, Alex Auvolat, Soumith Chintala, Timothée Lacroix, Zeming Lin, Florian Richoux, and Nicolas Usunier · 2016
Later among the works it cites.
Value Iteration Networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Towards Adapting Deep Visuomotor Representations from Simulated to Real Environments
Eric Tzeng, Coline Devin, Judy Hoffman, Chelsea Finn, Xingchao Peng, Sergey Levine, Kate Saenko, and Trevor Darrell · 2016
Later among the works it cites.
Deep Reinforcement Learning with Double Q-Learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
Strategic Attentive Writer for Learning Macro-Actions
Alexander Vezhnevets, Volodymyr Mnih, Simon Osindero, Alex Graves, Oriol Vinyals, John Agapiou, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Dueling Network Architectures for Deep Reinforcement Learning
Ziyu Wang, Nando de Freitas, and Marc Lanctot · 2016
Later among the works it cites.
The Option-Critic Architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Closest in time.
An Actor-Critic Algorithm for Sequence Prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Closest in time.
A Distributional Perspective on Reinforcement Learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Closest in time.
Recurrent Environment Simulators
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Closest in time.
Learning to Perform Physics Experiments via Deep Reinforcement Learning
Misha Denil, Pulkit Agrawal, Tejas D Kulkarni, Tom Erez, Peter Battaglia, and Nando de Freitas · 2017
Closest in time.
The Reactor: A Sample-Efficient Actor-Critic Architecture
Audrunas Gruslys, Mohammad Gheshlaghi Azar, Marc G Bellemare, and Rémi Munos · 2017
Closest in time.
Emergence of Locomotion Behaviours in Rich Environments
Nicolas Heess, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, Ali Eslami, Martin Riedmiller, et al · 2017
Closest in time.
Learning from Demonstrations for Real World Reinforcement Learning
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John Agapiou, et al · 2017
Closest in time.
Reinforcement Learning with Unsupervised Auxiliary Tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2017
Closest in time.
Schema Networks: Zero-Shot Transfer with a Generative Causal Model of Intuitive Physics
Ken Kansky, Tom Silver, David A Mély, Mohamed Eldawy, Miguel Lázaro-Gredilla, Xinghua Lou, Nimrod Dorfman, Szymon Sidor, Scott Phoenix, and Dileep George · 2017
Closest in time.
Multi-Agent Reinforcement Learning in Sequential Social Dilemmas
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Closest in time.
Learning to Optimize
Ke Li and Jitendra Malik · 2017
Closest in time.
Deep Reinforcement Learning: An Overview
Yuxi Li · 2017
Closest in time.
Discrete Sequential Prediction of Continuous Actions for Deep RL
Luke Metz, Julian Ibarz, Navdeep Jaitly, and James Davidson · 2017
Closest in time.
Learning to Navigate in Complex Environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andy Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2017
Closest in time.
Bridging the Gap Between Value and Policy Based Reinforcement Learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Closest in time.
Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine · 2017
Closest in time.
PGQ: Combining Policy Gradient and Q-Learning
Brendan O’Donoghue, Rémi Munos, Koray Kavukcuoglu, and Volodymyr Mnih · 2017
Closest in time.
Neural Map: Structured Memory for Deep Reinforcement Learning
Emilio Parisotto and Ruslan Salakhutdinov · 2017
Closest in time.
Learning Model-Based Planning from Scratch
Razvan Pascanu, Yujia Li, Oriol Vinyals, Nicolas Heess, Lars Buesing, Sebastien Racanière, David Reichert, Théophane Weber, Daan Wierstra, and Peter Battaglia · 2017
Closest in time.
Curiosity-Driven Exploration by Self-supervised Prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Closest in time.
Peng Peng, Ying Wen, Yaodong Yang, Quan Yuan, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Closest in time.
Neural Episodic Control
Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adrià Puigdomènech, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell · 2017
Closest in time.
Sim-to-Real Robot Learning from Pixels with Progressive Nets
Andrei A Rusu, Matej Vecerik, Thomas Rothörl, Nicolas Heess, Razvan Pascanu, and Raia Hadsell · 2017
Closest in time.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Tim Salimans, Jonathan Ho, Xi Chen, and Ilya Sutskever · 2017
Closest in time.
Third Person Imitation Learning
Bradley C Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
Closest in time.
#Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2017
Closest in time.
Distral: Robust Multitask Reinforcement Learning
Yee Whye Teh, Victor Bapst, Wojciech Marian Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Closest in time.
A Deep Hierarchical Approach to Lifelong Learning in Minecraft
Chen Tessler, Shahar Givony, Tom Zahavy, Daniel J Mankowitz, and Shie Mannor · 2017
Closest in time.
Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks
Nicolas Usunier, Gabriel Synnaeve, Zeming Lin, and Soumith Chintala · 2017
Closest in time.
FeUdal Networks for Hierarchical Reinforcement Learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Closest in time.
StarCraft II: A New Challenge for Reinforcement Learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al · 2017
Closest in time.
Imagination-Augmented Agents for Deep Reinforcement Learning
Théophane Weber, Sébastien Racanière, David P Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Closest in time.
Target-Driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning
Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi · 2017
Closest in time.
Neural Architecture Search with Reinforcement Learning
Barret Zoph and Quoc V Le · 2017
Closest in time.