Fetching the paper…
Reading the bibliography…
Transformer-based models have demonstrated remarkable in-context learning capabilities, prompting extensive research into its underlying mechanisms.
Iterative berechung der reziproken matrix
Schulz, G · 1933
Earlier work this paper cites.
Iterative methods of matrix inversion
Ogden, H. C · 1969
Earlier work this paper cites.
Matrix perturbation theory
Stewart, G. W · 1990
Earlier work this paper cites.
An improved newton iteration for the generalized inverse of a matrix, with applications
Pan, V · 1991
Earlier work this paper cites.
Convex optimization
Boyd, S. P · 2004
Earlier work this paper cites.
A family of iterative methods for computing the approximate inverse of a square matrix and inner inverse of a non-square matrix
Li, W · 2010
Earlier work this paper cites.
Large-scale mimo detection for 3gpp lte: Algorithms and fpga implementations
Wu, M · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization of self-concordant empirical loss
Zhang, Y · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Lectures on convex optimization
Nesterov, Y · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2020
Cited alongside, same era.
Composite convex optimization with global and local inexact oracles
Sun, T · 2020
Cited alongside, same era.
Jurassic-1: Technical details and evaluation
Lieber, O · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Nye, M · 2021
Cited alongside, same era.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E · 2022
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K · 2023
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Bai, Y · 2023
Later among the works it cites.
Transformers implement functional gradient descent to learn non-linear functions in context
Cheng, X · 2023
Later among the works it cites.
Transformers learn higher-order optimization methods for in-context learning: A study with linear models
Fu, D · 2023
Later among the works it cites.
Looped transformers as programmable computers
Giannou, A · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gpt-neox-20b: An open-source autoregressive language model
Black, S · 2022
Cited alongside, same era.
Language models show human-like content effects on reasoning
Dasgupta, I · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S · 2022
Cited alongside, same era.
Transformers in vision: A survey
Khan, S · 2022
Cited alongside, same era.
Transformers learn in-context by gradient descent
von Oswald, J · 2022
Cited alongside, same era.
Teaching algorithmic reasoning via in-context learning
Zhou, H · 2022
Cited alongside, same era.
Guo, T · 2023
Later among the works it cites.
In-context convergence of transformers
Huang, Y · 2023
Later among the works it cites.
Transformers as algorithms: Generalization and stability in in-context learning
Li, Y · 2023
Later among the works it cites.
Mahankali, A · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Zhang, R · 2023
Later among the works it cites.