Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms update an agent's parameters according to one of several possible rules, discovered manually through years of research.
Evolutionary principles in self-referential learning
J. Schmidhuber · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Y. Bengio, S. Bengio, and J. Cloutier · 1990
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
R. S. Sutton · 1992
Earlier work this paper cites.
A ‘self-referential’weight matrix
J. Schmidhuber · 1993
Earlier work this paper cites.
Learning one more thing
S. Thrun and T. M. Mitchell · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Automatic discovery of ranking formulas for playing with multi-armed bandits
F. Maes, L. Wehenkel, and D. Ernst · 2011
Earlier work this paper cites.
Policy search in a space of simple closed-form formulas: towards interpretability of reinforcement learning
F. Maes, R. Fonteneau, L. Wehenkel, and D. Ernst · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
A neuroevolution approach to general atari game playing
M. Hausknecht, J. Lehman, R. Miikkulainen, and P. Stone · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap · 2016
Earlier work this paper cites.
Matching networks for one shot learning
O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al · 2016
Earlier work this paper cites.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
N. P. Jouppi, C. Young, N. Patil, D. A. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, et al · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, and S. Wanderman-Milne · 2018
Cited alongside, same era.
J. Clune · 2019
Later among the works it cites.
Meta-learning with warped gradient descent
S. Flennerhag, A. A. Rusu, R. Pascanu, H. Yin, and R. Hadsell · 2019
Later among the works it cites.
Backpropamine: training self-modifying neural networks with differentiable neuromodulated plasticity
T. Miconi, A. Rawal, J. Clune, and K. O. Stanley · 2019
Later among the works it cites.
Badger: Learning to (learn [learning algorithms] through multi-agent communication)
M. Rosa, O. Afanasjeva, S. Andersson, J. Davidson, N. Guttenberg, P. Hlubuček, M. Poliak, J. Vítku, and J. Feyereisl · 2019
Later among the works it cites.
Discovery of useful questions as auxiliary tasks
V. Veeriah, M. Hessel, Z. Xu, J. Rajendran, R. L. Lewis, J. Oh, H. P. van Hasselt, D. Silver, and S. Singh · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Distributional reinforcement learning with quantile regression
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Cited alongside, same era.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
C. Finn and S. Levine · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Cited alongside, same era.
Evolved policy gradients
R. Houthooft, Y. Chen, P. Isola, B. Stadie, F. Wolski, O. J. Ho, and P. Abbeel · 2018
Cited alongside, same era.
Differentiable plasticity: training plastic neural networks with backpropagation
T. Miconi, J. Clune, and K. O. Stanley · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
A. Nichol, J. Achiam, and J. Schulman · 2018
Cited alongside, same era.
Beyond exponentially discounted sum: Automatic learning of return function
Y. Wang, Q. Ye, and T.-Y. Liu · 2019
Later among the works it cites.
RLax: Reinforcement Learning in JAX, 2020
D. Budden, M. Hessel, J. Quan, S. Kapturowski, K. Baumli, S. Bhupatiraju, A. Guy, and M. King · 2020
Closest in time.
Finding online neural update rules by learning to remember
K. Gregor · 2020
Closest in time.
Haiku: Sonnet for JAX, 2020
T. Hennigan, T. Cai, T. Norman, and I. Babuschkin · 2020
Closest in time.
Optax: composable gradient transformation and optimisation, in jax!, 2020
M. Hessel, D. Budden, F. Viola, M. Rosca, E. Sezener, and T. Hennigan · 2020
Closest in time.
Improving generalization in meta reinforcement learning using learned objectives
L. Kirsch, S. van Steenkiste, and J. Schmidhuber · 2020
Closest in time.
Behaviour suite for reinforcement learning
I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepezvari, S. Singh, et al · 2020
Closest in time.
Meta-gradient reinforcement learning with an objective discovered online
Z. Xu, H. van Hasselt, M. Hessel, J. Oh, S. Singh, and D. Silver · 2020
Closest in time.
A self-tuning actor-critic algorithm
T. Zahavy, Z. Xu, V. Veeriah, M. Hessel, J. Oh, H. van Hasselt, D. Silver, and S. Singh · 2020
Closest in time.
What can learned intrinsic rewards capture?
Z. Zheng, J. Oh, M. Hessel, Z. Xu, M. Kroiss, H. van Hasselt, D. Silver, and S. Singh · 2020
Closest in time.
Online meta-critic learning for off-policy actor-critic methods
W. Zhou, Y. Li, Y. Yang, H. Wang, and T. M. Hospedales · 2020
Closest in time.