Fetching the paper…
Reading the bibliography…
We reformulate the problem of supervised learning as learning to learn with two nested loops (i.e.
Variable kernel estimates of multivariate densities
Leo Breiman, William Meisel, and Edward Purcell · 1977
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Using fast weights to deblur old memories
Geoffrey E Hinton and David C Plaut · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
The nadaraya-watson kernel regression function estimator
Hermanus Josephus Bierens · 1988
Earlier work this paper cites.
Learning a synaptic learning rule
Yoshua Bengio, Samy Bengio, and Jocelyn Cloutier · 1990
Earlier work this paper cites.
Local learning algorithms
Léon Bottou and Vladimir Vapnik · 1992
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Learning by transduction
A. Gammerman, V. Vovk, and V. Vapnik · 1998
Earlier work this paper cites.
Learning to learn: Introduction and overview
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Weighted nadaraya–watson regression estimation
Zongwu Cai · 2001
Earlier work this paper cites.
Learning to classify text using support vector machines , volume 668
Thorsten Joachims · 2002
Earlier work this paper cites.
Large scale transductive svms
Ronan Collobert, Fabian Sinz, Jason Weston, Léon Bottou, and Thorsten Joachims · 2006
Earlier work this paper cites.
Tent: Fully test-time adaptation by entropy minimization
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2006
Earlier work this paper cites.
Gaussian processes for machine learning , volume 2
Christopher KI Williams and Carl Edward Rasmussen · 2006
Earlier work this paper cites.
Svm-knn: Discriminative nearest neighbor classification for visual category recognition
Hao Zhang, Alexander C Berg, Michael Maire, and Jitendra Malik · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Using fast weights to improve persistent contrastive divergence
Tijmen Tieleman and Geoffrey Hinton · 2009
Earlier work this paper cites.
Online domain adaptation of a pre-trained cascade of classifiers
Vidit Jain and Erik Learned-Miller · 2011
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Cited alongside, same era.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Cited alongside, same era.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros · 2016
Cited alongside, same era.
A tutorial on kernel density estimation and recent advances
Yen-Chi Chen · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Going beyond linear transformers with recurrent fast weight programmers
Kazuki Irie, Imanol Schlag, Róbert Csordás, and Jürgen Schmidhuber · 2021
Later among the works it cites.
Meta learning backpropagation and improving it
Louis Kirsch and Jürgen Schmidhuber · 2021
Later among the works it cites.
Ttt++: When does self-supervised test-time training fail or thrive?
Yuejiang Liu, Parth Kothari, Bastien van Delft, Baptiste Bellot-Gurlet, Taylor Mordan, and Alexandre Alahi · 2021
Later among the works it cites.
Soft: Softmax-free transformer with linear complexity
Jiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu, Weiguo Gao, Chunjing Xu, Tao Xiang, and Li Zhang · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Later among the works it cites.
Online learning of unknown dynamics for model-based controllers in legged locomotion
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Cited alongside, same era.
Meta-learning update rules for unsupervised representation learning
Luke Metz, Niru Maheswaranathan, Brian Cheung, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Online model distillation for efficient video inference
Ravi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan, and Kayvon Fatahalian · 2018
Cited alongside, same era.
“zero-shot” super-resolution using deep internal learning
Assaf Shocher, Nadav Cohen, and Michal Irani · 2018
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Cited alongside, same era.
Yu Sun, Wyatt L Ubellacker, Wen-Loong Ma, Xiang Zhang, Changhao Wang, Noel V Csomay-Shanklin, Masayoshi Tomizuka, Koushil Sreenath, and Aaron D Ames · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Later among the works it cites.
Better plain vit baselines for imagenet-1k
Lucas Beyer, Xiaohua Zhai, and Alexander Kolesnikov · 2022
Later among the works it cites.
Meta-learning fast weight language models
Kevin Clark, Kelvin Guu, Ming-Wei Chang, Panupong Pasupat, Geoffrey Hinton, and Mohammad Norouzi · 2022
Later among the works it cites.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2022
Later among the works it cites.
Test-time training with masked autoencoders
Yossi Gandelsman, Yu Sun, Xinlei Chen, and Alexei A. Efros · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Geoffrey Hinton · 2022
Later among the works it cites.
Mystyle: A personalized generative prior
Yotam Nitzan, Kfir Aberman, Qiurui He, Orly Liba, Michal Yarom, Yossi Gandelsman, Inbar Mosseri, Yael Pritch, and Daniel Cohen-Or · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Closest in time.
Test-time training on nearest neighbors for large language models
Moritz Hardt and Yu Sun · 2023
Closest in time.
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2023
Closest in time.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Closest in time.
Transformers as support vector machines
Davoud Ataee Tarzanagh, Yingcong Li, Christos Thrampoulidis, and Samet Oymak · 2023
Closest in time.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Closest in time.
Test-time training on video streams
Renhao Wang, Yu Sun, Yossi Gandelsman, Xinlei Chen, Alexei A Efros, and Xiaolong Wang · 2023
Closest in time.