Fetching the paper…
Reading the bibliography…
This work identifies a simple pre-training mechanism that leads to representations exhibiting better continual and transfer learning.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
M. McCloskey and N. J. Cohen · 1989
Earlier work this paper cites.
Learning a synaptic learning rule
Y. Bengio, S. Bengio, and J. Cloutier · 1990
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
R. M. French · 1999
Earlier work this paper cites.
Learning to learn using gradient descent
S. Hochreiter, A. S. Younger, and P. R. Conwell · 2001
Earlier work this paper cites.
Robustness and evolution: concepts, insights and challenges from a developmental model system
M.-A. Félix and A. Wagner · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Emergent network structure, evolvable robustness, and nonlinear effects of point mutations in an artificial genome model
T. Rohlf and C. R. Winkler · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
D. Ulyanov, A. Vedaldi, and V. Lempitsky · 2016
Cited alongside, same era.
Matching networks for one shot learning
O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2017
Cited alongside, same era.
Meta-learning by the baldwin effect
C. Fernando, J. Sygnowski, S. Osindero, J. Wang, T. Schaul, D. Teplyashin, P. Sprechmann, A. Pritzel, and A. Rusu · 2018
Cited alongside, same era.
Fix your classifier: The marginal value of training the last weight layer
The impact of reinitialization on generalization in convolutional neural networks
I. Alabdulmohsin, H. Maennel, and D. Keysers · 2021
Later among the works it cites.
Meta-baseline: Exploring simple meta-learning for few-shot learning
Y. Chen, Z. Liu, H. Xu, T. Darrell, and X. Wang · 2021
Later among the works it cites.
Continual backprop: Stochastic gradient descent with persistent randomness
S. Dohare, A. R. Mahmood, and R. S. Sutton · 2021
Later among the works it cites.
Knowledge evolution in neural networks
A. Taha, A. Shrivastava, and L. S. Davis · 2021
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
A. Kumar, A. Raghunathan, R. M. Jones, T. Ma, and P. Liang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Hoffer, I. Hubara, and D. Soudry · 2018
Cited alongside, same era.
A practitioners’ guide to transfer learning for text classification using convolutional neural networks
T. Semwal, P. Yenigalla, G. Mathur, and S. B. Nair · 2018
Cited alongside, same era.
Retraining: A simple way to improve the ensemble accuracy of deep neural networks for image classification
K. Zhao, T. Matsukawa, and E. Suzuki · 2018
Cited alongside, same era.
Meta-learning representations for continual learning
K. Javed and M. White · 2019
Cited alongside, same era.
S. Beaulieu, L. Frati, T. Miconi, J. Lehman, K. O. Stanley, J. Clune, and N. Cheney · 2020
Cited alongside, same era.
The early phase of neural network training
J. Frankle, D. J. Schwab, and A. S. Morcos · 2020
Cited alongside, same era.
Rifle: Backpropagation in depth for deep transfer learning through re-initializing the fully-connected layer
X. Li, H. Xiong, H. An, C.-Z. Xu, and D. Dou · 2020
Cited alongside, same era.
The primacy bias in deep reinforcement learning
E. Nikishin, M. Schwarzer, P. D’Oro, P.-L. Bacon, and A. Courville · 2022
Later among the works it cites.
When does re-initialization work?
S. Zaidi, T. Berariu, H. Kim, J. Bornschein, C. Clopath, Y. W. Teh, and R. Pascanu · 2022
Later among the works it cites.
Are all layers created equal?
C. Zhang, S. Bengio, and Y. Singer · 2022
Later among the works it cites.
Fortuitous forgetting in connectionist networks
H. Zhou, A. Vani, H. Larochelle, and A. Courville · 2022
Later among the works it cites.
Improving language plasticity via pretraining with active forgetting
Y. Chen, K. Marchisio, R. Raileanu, D. I. Adelani, P. Stenetorp, S. Riedel, and M. Artetxe · 2023
Closest in time.
Omnimage: Evolving 1k image cliques for few-shot learning
L. Frati, N. Traft, and N. Cheney · 2023
Closest in time.
Understanding plasticity in neural networks
C. Lyle, Z. Zheng, E. Nikishin, B. A. Pires, R. Pascanu, and W. Dabney · 2023
Closest in time.
Learn, unlearn and relearn: An online learning paradigm for deep neural networks
V. R. T. Ramkumar, E. Arani, and B. Zonooz · 2023
Closest in time.