Fetching the paper…
Reading the bibliography…
There is a long history of using meta learning as representation learning, specifically for determining the relevance of inputs.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton, R. S. (1992) · 1992
Earlier work this paper cites.
Temporal-difference methods and Markov models
Barnard, E. (1993) · 1993
Earlier work this paper cites.
A model of cerebellar metaplasticity
Schweighofer, N. and Arbib, M. A. (1998) · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
Schraudolph, N. N. (1999) · 1999
Earlier work this paper cites.
From the lexicon to expectations about kinds: A role for associative learning
Colunga, E. and Smith, L. B. (2005) · 2005
Earlier work this paper cites.
What is intrinsic motivation? A typology of computational approaches
Oudeyer, P.-Y. and Kaplan, F. (2009) · 2009
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Earlier work this paper cites.
Adaptive Step-Size for Online Temporal Difference Learning
Dabney, W. and Barto, A. G. (2012) · 2012
Cited alongside, same era.
Tuning-free step-size adaptation
Mahmood, A. R., Sutton, R. S., Degris, T., and Pilarski, P. M. (2012) · 2012
Cited alongside, same era.
Dynamic switching and real-time machine learning for improved human control of assistive biomedical robots
Pilarski, P. M., Dawson, M. R., Degris, T., Carey, J. P., and Sutton, R. S. (2012) · 2012
Cited alongside, same era.
Inferring relevance in a changing world
Wilson, R. C. and Niv, Y. (2012) · 2012
Cited alongside, same era.
Representation Search through Generate and Test
Mahmood, A. R. and Sutton, R. S. (2013) · 2013
Cited alongside, same era.
Adaptive artificial limbs: A real-time approach to prediction and anticipation
Pilarski, P. M., Dawson, M. R., Degris, T., Carey, J. P., Chan, K. M., Hebert, J. S., and Sutton, R. S. (2013) · 2013
True online TD (lambda)
Seijen, H. and Sutton, R. (2014) · 2014
Later among the works it cites.
Machine learning and unlearning to autonomously switch between the functions of a myoelectric arm
Edwards, A. L., Hebert, J. S., and Pilarski, P. M. (2016) · 2016
Later among the works it cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Later among the works it cites.
Every step you take: Vectorized adaptive step sizes for temporal difference learning
Kearney, A., Veeriah, V., Travnik, J., Sutton, R. S., and Pilarski, P. M. (2017) · 2017
Later among the works it cites.
Representing high-dimensional data to intelligent prostheses and other wearable assistive robots: A first comparison of tile coding and selective Kanerva coding
Travnik, J. B. and Pilarski, P. M. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adaptive Step-Sizes for Reinforcement Learning
Dabney, W. C. (2014) · 2014
Cited alongside, same era.
Development of the Bento Arm: An improved robotic arm for myoelectric training and research
Dawson, M. R., Sherstan, C., Carey, J. P., Hebert, J. S., and Pilarski, P. M. (2014) · 2014
Cited alongside, same era.
Meta-Gradient Reinforcement Learning
Xu, Z., van Hasselt, H., and Silver, D. (2018) · 2018
Later among the works it cites.
Metatrace: Online Step-size Tuning by Meta-gradient Descent for Reinforcement Learning Control
Young, K., Wang, B., and Taylor, M. E. (2018) · 2018
Later among the works it cites.