Fetching the paper…
Reading the bibliography…
We study gradient flow on the exponential loss for a classification problem with a one-layer softmax attention model, where the key and query weight matrices are trained separately.
Language models are few-shot learners
Brown, T · 1901
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K · 1906
Earlier work this paper cites.
Characterization of the subdifferential of some matrix norms
Watson, G. A · 1992
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S · 2006
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2010
Earlier work this paper cites.
A unifying view on implicit bias in training linear neural networks
Yun, C · 2010
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A · 2012
Earlier work this paper cites.
Li, Z · 2012
Earlier work this paper cites.
Approximate kkt points and a proximity measure for termination
Dutta, J · 2013
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Stochastic gradient/mirror descent: Minimax optimality and implicit regularization
Azizan, N · 2018
Earlier work this paper cites.
Gradient descent aligns the layers of deep linear networks
Ji, Z · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Soudry, D · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Arora, S · 2019
Earlier work this paper cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gidel, G · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Vaskevicius, T · 2019
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ji, Z · 2020
Cited alongside, same era.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Moroshko, E · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Woodworth, B · 2020
Cited alongside, same era.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Azulay, S · 2021
Cited alongside, same era.
On the optimization and generalization of multi-head attention
Deora, P · 2023
Later among the works it cites.
Understanding implicit regularization in over-parameterized single index model
Fan, J · 2023
Later among the works it cites.
In-context convergence of transformers
Huang, Y · 2023
Later among the works it cites.
On the role of attention in prompt-tuning
Oymak, S · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z · 2021
Cited alongside, same era.
Fast margin maximization via dual acceleration
Ji, Z · 2021
Cited alongside, same era.
Characterizing the implicit bias via a primal-dual analysis
Ji, Z · 2021
Cited alongside, same era.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Stöger, D · 2021
Cited alongside, same era.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L · 2021
Cited alongside, same era.
A unifying view on implicit bias in training linear neural networks
Yun, C · 2021
Cited alongside, same era.
Wind, J. S · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Zhang, R · 2023
Later among the works it cites.
Implicit regularization leads to benign overfitting for sparse linear regression
Zhou, M · 2023
Later among the works it cites.
Chen, S · 2024
Closest in time.
(s) gd over diagonal linear networks: Implicit bias, large stepsizes and edge of stability
Even, M · 2024
Closest in time.
Transformers provably learn feature-position correlations in masked image modeling
Huang, Y · 2024
Closest in time.
How transformers learn causal structure with gradient descent
Nichani, E · 2024
Closest in time.
Saddle-to-saddle dynamics in diagonal linear networks
Pesme, S · 2024
Closest in time.
Implicit bias of next-token prediction
Thrampoulidis, C · 2024
Closest in time.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Tian, Y · 2024
Closest in time.
Implicit bias and fast convergence rates for self-attention
Vasudeva, B · 2024
Closest in time.