Fetching the paper…
Reading the bibliography…
Attention-based neural networks such as transformers have demonstrated a remarkable ability to exhibit in-context learning (ICL): Given a short prompt sequence of tokens from an unseen task, they can formulate relevant per-token and next-token predictions without any parameter updates.
“The evaluation of the collision matrix”
Gian-Carlo Wick · 1950
Earlier work this paper cites.
“On a product of positive semidefinite matrices”
AR Meenakshi and C Rajian · 1999
Earlier work this paper cites.
“The matrix cookbook”
Kaare Petersen and Michael Pedersen · 2008
Earlier work this paper cites.
“An Isserlis’ theorem for mixed Gaussian variables: Application to the auto-bispectral density”
JV Michalowicz, JM Nichols, F Bucholtz and CC Olson · 2009
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Implicit regularization in matrix factorization”
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur and Nati Srebro · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“On the optimization of deep networks: Implicit acceleration by overparameterization”
Sanjeev Arora, Nadav Cohen and Elad Hazan · 2018
Earlier work this paper cites.
“Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced”
Simon Du, Wei Hu and Jason Lee · 2018
Earlier work this paper cites.
“Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations”
Yuanzhi Li, Tengyu Ma and Hongyang Zhang · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training”
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“Implicit regularization in deep matrix factorization”
Sanjeev Arora, Nadav Cohen, Wei Hu and Yuping Luo · 2019
Earlier work this paper cites.
“Nonconvex optimization meets low-rank matrix factorization: An overview”
Yuejie Chi, Yue Lu and Yuxin Chen · 2019
Earlier work this paper cites.
“Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc. Le and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
“Universal Transformers”, 2019
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit and Łukasz Kaiser · 2019
Earlier work this paper cites.
“On the turing completeness of modern neural network architectures”
Jorge Pérez, Javier Marinković and Pablo Barceló · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“Are transformers universal approximators of sequence-to-sequence functions?”
Chulhee Yun, Srinadh Bhojanapalli, Ankit Rawat, Sashank Reddi and Sanjiv Kumar · 2019
Cited alongside, same era.
“On implicit regularization: Morse functions and applications to matrix factorization”
Mohamed Belabbas · 2020
Cited alongside, same era.
“On the computational power of transformers and its implications in sequence modeling”
Satwik Bhattamishra, Arkil Patel and Navin Goyal · 2020
Cited alongside, same era.
Zhiyuan Li, Yuping Luo and Kaifeng Lyu · 2020
Cited alongside, same era.
“Transformers: State-of-the-art natural language processing”
“Transformers learn in-context by gradient descent”
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov and Max Vladymyrov · 2022
Later among the works it cites.
“A Mechanism for Sample-Efficient In-Context Learning for Sparse Retrieval Tasks”
Jacob Abernethy, Alekh Agarwal, Teodor. Marinov and Manfred. Warmuth · 2023
Closest in time.
“Transformers learn to implement preconditioned gradient descent for in-context learning”
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand and Suvrit Sra · 2023
Closest in time.
“In-Context Learning through the Bayesian Prism”
Kabir Ahuja, Madhur Panwar and Navin Goyal · 2023
Closest in time.
“A Closer Look at In-Context Learning under Distribution Shifts”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf and Morgan Funtowicz · 2020
Cited alongside, same era.
“O (n) connections are expressive enough: Universal approximability of sparse transformers”
Chulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Rawat, Sashank Reddi and Sanjiv Kumar · 2020
Cited alongside, same era.
“On the implicit bias of initialization shape: Beyond infinitesimal mirror descent”
Shahar Azulay, Edward Moroshko, Mor Nacson, Blake Woodworth, Nathan Srebro, Amir Globerson and Daniel Soudry · 2021
Cited alongside, same era.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit and Neil Houlsby · 2021
Cited alongside, same era.
“On the expressive power of self-attention matrices”
Valerii Likhosherstov, Krzysztof Choromanski and Adrian Weller · 2021
Cited alongside, same era.
“An explanation of in-context learning as implicit bayesian inference”
Sang Xie, Aditi Raghunathan, Percy Liang and Tengyu Ma · 2021
Cited alongside, same era.
“What learning algorithm is in-context learning? Investigations with linear models”
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma and Denny Zhou · 2022
Cited alongside, same era.
“Exploring Length Generalization in Large Language Models”
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer and Behnam Neyshabur · 2022
Cited alongside, same era.
Kartik Ahuja and David Lopez-Paz · 2023
Closest in time.
“Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection”
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong and Song Mei · 2023
Closest in time.
“In-Context Learning of Large Language Models Explained as Kernel Regression”, 2023
Chi Han, Ziqi Wang, Han Zhao and Heng Ji · 2023
Closest in time.
“Understanding incremental learning of gradient descent: A fine-grained analysis of matrix sensing”
Jikai Jin, Zhiyuan Li, Kaifeng Lyu, Simon Du and Jason Lee · 2023
Closest in time.
“The Closeness of In-Context Learning and Weight Shifting for Softmax Regression”
Shuai Li, Zhao Song, Yu Xia, Tong Yu and Tianyi Zhou · 2023
Closest in time.
“Transformers as Algorithms: Generalization and Stability in In-context Learning”
Yingcong Li, M Ildiz, Dimitris Papailiopoulos and Samet Oymak · 2023
Closest in time.
“How do transformers learn topic structure: Towards a mechanistic understanding”
Yuchen Li, Yuanzhi Li and Andrej Risteski · 2023
Closest in time.
“Transformers Learn Shortcuts to Automata”
Bingbin Liu, Jordan. Ash, Surbhi Goel, Akshay Krishnamurthy and Cyril Zhang · 2023
Closest in time.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.
Mahdi Soltanolkotabi, Dominik Stöger and Changzhi Xie · 2023
Closest in time.
“Mimetic Initialization of Self-Attention Layers”
Asher Trockman and J Kolter · 2023
Closest in time.
Xinyi Wang, Wanrong Zhu and William Wang · 2023
Closest in time.
Yufeng Zhang, Fengzhuo Zhang, Zhuoran Yang and Zhaoran Wang · 2023
Closest in time.