Fetching the paper…
Reading the bibliography…
We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs).
Transfer in deep reinforcement learning using successor features and generalised policy improvement
A. Barreto, D. Borsa, J. Quan, T. Schaul, D. Silver, M. Hessel, D. Mankowitz, A. Zídek, and R. Munos · 1901
Earlier work this paper cites.
Reinforcement today
B. F. Skinner · 1958
Earlier work this paper cites.
Distribution of eigenvalues or some sets of random matrices
V. A. Marŏenko and L. A. Pastur · 1967
Earlier work this paper cites.
The variance of discounted Markov decision processes
M. J. Sobel · 1982
Earlier work this paper cites.
A dual back-propagation scheme for scalar reinforcement learning
P. W. Munro · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Dynamic Error Propagation Networks
A. J. Robinson · 1989
Earlier work this paper cites.
Dynamic reinforcement driven error propagation networks with application to game playing
T. Robinson and F. Fallside · 1989
Earlier work this paper cites.
Learning from Delayed Rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Implementierung und Anwendung eines ‘neuronalen’ Echtzeit-Lernalgorithmus für reaktive Umgebungen
S. Hochreiter · 1990
Earlier work this paper cites.
Markov decision processes
M. L. Puterman · 1990
Earlier work this paper cites.
Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments
J. Schmidhuber · 1990
Earlier work this paper cites.
Solving h h -horizon, stationary Markov decision problems in time proportional to log ( h ) \log(h)
P. Tseng · 1990
Earlier work this paper cites.
A menu of designs for reinforcement learning over time
P. J. Werbos · 1990
Earlier work this paper cites.
An analysis of stochastic shortest path problems
D. P. Bertsekas and J. N. Tsitsiklis · 1991
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Reinforcement learning in markovian and non-markovian environments
J. Schmidhuber · 1991
Earlier work this paper cites.
The convergence of TD( λ \lambda ) for general λ \lambda
P. Dayan · 1992
Earlier work this paper cites.
A note on continuity of fixed points
M. Kwiecinski · 1992
Earlier work this paper cites.
Q-Learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Reinforcement Learning for Robots Using Neural Networks
L. Lin · 1993
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
A. W. Moore and C. G. Atkeson · 1993
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
T. Jaakkola, M. I. Jordan, and S. P. Singh · 1994
Earlier work this paper cites.
When the best move isn’t optimal: q q -learning with exploration
G. H. John · 1994
Earlier work this paper cites.
On-line q q -learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and q q -learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1995
Earlier work this paper cites.
Neuro-dynamic programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Continuous dependence of attractors of iterated function systems
J. Jachymski · 1996
Earlier work this paper cites.
Incremental multi-step q q -learning
J. Peng and R. J. Williams · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
S. P. Singh and R. S. Sutton · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
V. S. Borkar · 1997
Earlier work this paper cites.
Recurrent neural net learning and vanishing gradient
S. Hochreiter · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
LSTM can solve hard long time lag problems
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Continuous dependence on parameters of the fixed point set for some set-valued operators
E. Kirr and A. Petrusel · 1997
Earlier work this paper cites.
Stochastic and shortest path games: theory and algorithms
S. D. Patek · 1997
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
S. Hochreiter · 1998
Earlier work this paper cites.
Comparing value-function estimation algorithms in undiscounted problems
F. Beleznay, T. Grobler, and C. Szepesvári · 1999
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
F. A. Gers, J. Schmidhuber, and F. Cummins · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. J. Russell · 1999
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
S. Schaal · 1999
Earlier work this paper cites.
Recurrent nets that time and count
F. A. Gers and J. Schmidhuber · 2000
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
F. A. Gers, J. Schmidhuber, and F. Cummins · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
S. Singh, T. Jaakkola, M. Littman, and C. Szepesvári · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Learning to learn using gradient descent
S. Hochreiter, A. Steven Younger, and Peter R. Conwell · 2001
Earlier work this paper cites.
Independent Component Analysis
A. Hyvärinen, J. Karhunen, and E. Oja · 2001
Earlier work this paper cites.
Symmetries and model minimization in Markov decision processes
B. Ravindran and A. G. Barto · 2001
Cited alongside, same era.
Reinforcement learning with long short-term memory
B. Bakker · 2002
Cited alongside, same era.
A note on universality of the distribution of the largest eigenvalues in certain sample covariance matrices
A. Soshnikov · 2002
Cited alongside, same era.
Equivalence notions and model minimization in Markov decision processes
R. Givan, T. Dean, and M. Greig · 2003
Cited alongside, same era.
SMDP homomorphisms: An algebraic approach to abstraction in semi-Markov decision processes
B. Ravindran and A. G. Barto · 2003
Cited alongside, same era.
Potential-based shaping and q q -value initialization are equivalent
E. Wiewiora · 2003
Cited alongside, same era.
Deep learning in neural networks: An overview
J. Schmidhuber · 2015
Later among the works it cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel · 2015
Later among the works it cites.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Later among the works it cites.
Reward shaping with recurrent neural networks for speeding up on-line policy learning in spoken dialogue systems
P.-H. Su, D. Vandyke, M. Gasic, N. Mrksic, T.-H. Wen, and S. Young · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, and M. Lanctot · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Principled methods for advising reinforcement learning agents
E. Wiewiora, G. Cottrell, and C. Elkan · 2003
Cited alongside, same era.
Framewise phoneme classification with bidirectional LSTM and other neural network architectures
A. Graves and J. Schmidhuber · 2005
Cited alongside, same era.
Markov Decision Processes
M. L. Puterman · 2005
Cited alongside, same era.
Bandit based Monte-Carlo planning
L. Kocsis and C. Szepesvári · 2006
Cited alongside, same era.
Towards a unified theory of state abstraction for MDPs
L. Li, T. J. Walsh, and M. L. Littman · 2006
Cited alongside, same era.
Visual explanation of evidence in additive classifiers
B. Poulin, R. Eisner, D. Szafron, P. Lu, R. Greiner, D. S. Wishart, A. Fyshe, B. Pearcy, C. MacDonell, and J. Anvik · 2006
Cited alongside, same era.
Later among the works it cites.
Openai gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Later among the works it cites.
Continuity properties and sensitivity analysis of parameterized fixed points and approximate fixed points
Z. Feinstein · 2016
Later among the works it cites.
Learning and transfer of modulated locomotor controllers
N. Heess, G. Wayne, Y. Tassa, T. P. Lillicrap, M. A. Riedmiller, and D. Silver · 2016
Later among the works it cites.
On the analysis of complex backup strategies in Monte Carlo Tree Search
P. Khandelwal, E. Liebman, S. Niekum, and P. Stone · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Later among the works it cites.
Learning the variance of the reward-to-go
A. Tamar, D. DiCastro, and S. Mannor · 2016
Later among the works it cites.
Deep reinforcement learning with double q q -learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
Ergodic Markov processes and poisson equations (lecture notes)
A. Veretennikov · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. de Freitas · 2016
Later among the works it cites.
Top-down neural attention by excitation backprop
J. Zhang, Z. L. Lin, J. Brandt, X. Shen, and S. Sclaroff · 2016
Later among the works it cites.
Explaining recurrent neural network predictions in sentiment analysis
L. Arras, G. Montavon, K.-R. Müller, and W. Samek · 2017
Later among the works it cites.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. Hunt, T. Schaul, H. P. vanHasselt, and D. Silver · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Later among the works it cites.
BEGAN: boundary equilibrium generative adversarial networks
D. Berthelot, T. Schumm, and L. Metz · 2017
Later among the works it cites.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver · 2017
Later among the works it cites.
Two time-scale stochastic approximation with controlled Markov noise and off-policy temporal-difference learning
P. Karmakar and S. Bhatnagar · 2017
Later among the works it cites.
Time delays, competitive interdependence, and firm performance
J. Luoma, S. Ruutu, A. W. King, and H. Tikkanen · 2017
Later among the works it cites.
Explaining nonlinear classification decisions with deep Taylor decomposition
G. Montavon, S. Lapuschkin, A. Binder, W. Samek, and K.-R. Müller · 2017
Later among the works it cites.
Methods for interpreting and understanding deep neural networks
G. Montavon, W. Samek, and K.-R. Müller · 2017
Later among the works it cites.
The uncertainty Bellman equation and exploration
B. O’Donoghue, I. Osband, R. Munos, and V. Mnih · 2017
Later among the works it cites.
Mastering Chess and Shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. P. Lillicrap, K. Simonyan, and D. Hassabis · 2017
Later among the works it cites.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Later among the works it cites.
Playing hard exploration games by watching YouTube
Y. Aytar, T. Pfaff, D. Budden, T. Le Paine, Z. Wang, and N. de Freitas · 2018
Closest in time.
Forward-backward reinforcement learning
A. D. Edwards, L. Downs, and J. C. Davidson · 2018
Closest in time.
IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Closest in time.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg · 2018
Closest in time.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Closest in time.
Recall traces: Backtracking models for efficient reinforcement learning
A. Goyal, P. Brakel, W. Fedus, T. Lillicrap, S. Levine, H. Larochelle, and Y. Bengio · 2018
Closest in time.
World models
D. Ha and J. Schmidhuber · 2018
Closest in time.
Is multiagent deep reinforcement learning the answer or the question? A brief survey
P. Hernandez-Leal, B. Kartal, and M. E. Taylor · 2018
Closest in time.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver · 2018
Closest in time.
Optimizing agent behavior over long time scales by transporting value
C. Hung, T. Lillicrap, J. Abramson, Y. Wu, M. Mirza, F. Carnevale, A. Ahuja, and G. Wayne · 2018
Closest in time.
Sparse attentive backtracking: Temporal credit assignment through reminding
N. Ke, A. Goyal, O. Bilaniuk, J. Binas, M. Mozer, C. Pal, and Y. Bengio · 2018
Closest in time.
Bandit Algorithms
T. Lattimore and C. Szepesvá · 2018
Closest in time.
The machine learning reproducibility checklist, 2018
J. Pineau · 2018
Closest in time.
Observe and look further: Achieving consistent performance on Atari
T. Pohlen, B. Piot, T. Hester, M. G. Azar, D. Horgan, D. Budden, G. Barth-Maron, H. van Hasselt, J. Quan, M. Večerík, M. Hessel, R. Munos, and O. Pietquin · 2018
Closest in time.
Reward estimation for variance reduction in deep reinforcement learning
J. Romoff, A. Piché, P. Henderson, V. Francois-Lavet, and J. Pineau · 2018
Closest in time.
Reinforcement learning never worked, and ’deep’ only helped a bit
H. Sahni · 2018
Closest in time.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2018
Closest in time.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Closest in time.