Fetching the paper…
Reading the bibliography…
Some autoregressive models exhibit in-context learning capabilities: being able to learn as an input sequence is processed, without undergoing any parameter changes, and without being explicitly trained to do so.
The approximate arithmetical solution by finite differences of physical problems involving differential equations, with an application to the stresses in a masonry dam
Lewis Fry Rirchardson · 1911
Earlier work this paper cites.
The Organization of Behavior: A Neuropsychological Theory
Donald O. Hebb · 1949
Earlier work this paper cites.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
Jack Sherman and Winifred J. Morrison · 1950
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
R. E. Kalman · 1960
Earlier work this paper cites.
Adaptive switching circuits
Bernard Widrow and Marcian E. Hoff · 1960
Earlier work this paper cites.
Adaptive Control Processes: A Guided Tour
Richard Bellman · 1961
Earlier work this paper cites.
Chebyshev semi-iterative methods, successive overrelaxation iterative methods, and second order Richardson iterative methods: Part I
Gene H. Golub and Richard S. Varga · 1961
Earlier work this paper cites.
Exponential convergence of recursive least squares with exponential forgetting factor
Richard M Johnstone, C Richard Johnson Jr, Robert R Bitmead, and Brian DO Anderson · 1982
Earlier work this paper cites.
Introduction to the Theory of Neural Computation
John Hertz, Richard G. Palmer, and Anders S. Krogh · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: an alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Generalization in a linear perceptron in the presence of noise
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
The "wake-sleep" algorithm for unsupervised neural networks
Geoffrey E. Hinton, Peter Dayan, Brendan J. Frey, and Radford M. Neal · 1995
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects
Rajesh P. N. Rao and Dana H. Ballard · 1999
Earlier work this paper cites.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Richard H. R. Hahnloser, Rahul Sarpeshkar, Misha A. Mahowald, Rodney J. Douglas, and H. Sebastian Seung · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A. Steven Younger, and Peter R. Conwell · 2001
Earlier work this paper cites.
A distribution-free theory of nonparametric regression
László Györfi, Michael Kohler, Adam Krzyzak, and Harro Walk · 2002
Earlier work this paper cites.
Hierarchical Bayesian inference in the visual cortex
Tai Sing Lee and David Mumford · 2003
Earlier work this paper cites.
On intelligence
Jeff Hawkins and Sandra Blakeslee · 2004
Earlier work this paper cites.
A note on persistency of excitation
Jan C. Willems, Paolo Rapisarda, Ivan Markovsky, and Bart L. M. De Moor · 2005
Earlier work this paper cites.
A Fast Learning Algorithm for Deep Belief Nets
Geoffrey Hinton, Simon Osindero, and Yee Whye Teh · 2006
Earlier work this paper cites.
A free energy principle for the brain
Karl Friston, James Kilner, and Lee Harrison · 2006
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
Whatever next? Predictive brains, situated agents, and the future of cognitive science
Andy Clark · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Difference target propagation
Dong-Hyun Lee, Saizheng Zhang, Asja Fischer, and Yoshua Bengio · 2015
Earlier work this paper cites.
Adam: a method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
RL2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Cited alongside, same era.
OptNet: Differentiable optimization as a layer in neural networks
Brandon Amos and J. Zico Kolter · 2017
Cited alongside, same era.
An approximation of the error backpropagation algorithm in a predictive coding network with local Hebbian synaptic plasticity
James C. R. Whittington and Rafal Bogacz · 2017
Cited alongside, same era.
Principles of system identification: theory and practice
Arun K. Tangirala · 2018
Cited alongside, same era.
A general method for amortizing variational filtering
Joseph Marino, Milan Cvitkovic, and Yisong Yue · 2018
On the dangers of stochastic parrots: can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Later among the works it cites.
Deep declarative networks
Stephen Gould, Richard Hartley, and Dylan John Campbell · 2021
Later among the works it cites.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Lukas Gruber, Markus Holzleitner, Thomas Adler, David Kreil, Michael K. Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2021
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
Stephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang, Aaditya Singh, Pierre H. Richemond, Jay McClelland, and Felix Hill · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Predictive processing: a canonical cortical computation
Georg B. Keller and Thomas D. Mrsic-Flogel · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Meta-learners’ learning dynamics are unlike learners’
Neil C. Rabinowitz · 2019
Cited alongside, same era.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Cited alongside, same era.
Training neural networks with local error signals
Arild Nøkland and Lars Hiller Eidnes · 2019
Cited alongside, same era.
Later among the works it cites.
An explanation of in-context learning as implicit Bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Later among the works it cites.
What can transformers learn in-context? A case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S. Liang, and Gregory Valiant · 2022
Later among the works it cites.
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2022
Later among the works it cites.
General-purpose in-context learning by meta-learning transformers
Louis Kirsch, James Harrison, Jascha Sohl-Dickstein, and Luke Metz · 2022
Later among the works it cites.
Chess as a testbed for language model state tracking
Shubham Toshniwal, Sam Wiseman, Karen Livescu, and Kevin Gimpel · 2022
Later among the works it cites.
Beyond backpropagation: bilevel optimization through implicit differentiation and equilibrium propagation
Nicolas Zucchet and João Sacramento · 2022
Later among the works it cites.
The least-control principle for local learning at equilibrium
Alexander Meulemans, Nicolas Zucchet, Seijin Kobayashi, Johannes von Oswald, and João Sacramento · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Geoffrey Hinton · 2022
Later among the works it cites.
The transient nature of emergent in-context learning in transformers
Aaditya Singh, Stephanie Chan, Ted Moskovitz, Erin Grant, Andrew Saxe, and Felix Hill · 2023
Closest in time.
What learning algorithm is in-context learning? Investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Closest in time.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett · 2023
Closest in time.
Arvind Mahankali, Tatsunori B. Hashimoto, and Tengyu Ma · 2023
Closest in time.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Closest in time.
Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression
Allan Raventós, Mansheej Paul, Feng Chen, and Surya Ganguli · 2023
Closest in time.
CausalLM is not optimal for in-context learning
Nan Ding, Tomer Levinboim, Jialin Wu, Sebastian Goodman, and Radu Soricut · 2023
Closest in time.
Exploring the space of key-value-query models with intention
Marta Garnelo and Wojciech Marian Czarnecki · 2023
Closest in time.
Zoology: Measuring and improving recall in efficient language models
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Closest in time.
Kazuki Irie, Róbert Csordás, and Jürgen Schmidhuber · 2023
Closest in time.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Closest in time.
Energy transformer
Benjamin Hoover, Yuchen Liang, Bao Pham, Rameswar Panda, Hendrik Strobelt, Duen Horng Chau, Mohammed Zaki, and Dmitry Krotov · 2023
Closest in time.
Linear transformers are versatile in-context learners
Max Vladymyrov, Johannes von Oswald, Mark Sandler, and Rong Ge · 2024
Closest in time.
How well can transformers emulate in-context Newton’s method?
Angeliki Giannou, Liu Yang, Tianhao Wang, Dimitris Papailiopoulos, and Jason D Lee · 2024
Closest in time.
Linear attention is (maybe) all you need (to understand transformer optimization)
Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie, and Suvrit Sra · 2024
Closest in time.
Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, and Caglar Gulcehre · 2024
Closest in time.
Dual operating modes of in-context learning
Ziqian Lin and Kangwook Lee · 2024
Closest in time.