Fetching the paper…
Reading the bibliography…
We study the fundamental optimization principles of self-attention, the defining mechanism of transformers, by analyzing the implicit bias of gradient-based optimizers in training a self-attention layer with a linear decoder in binary classification.
Nuanced metrics for measuring unintended bias with real data for text classification
Borkan, D., Dixon, L., Sorensen, J., Thain, N., and Vasserman, L. (2019) · 1903
Earlier work this paper cites.
Harmless interpolation of noisy data in regression
Muthukumar, V., Vodrahalli, K., and Sahai, A. (2019) · 1903
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y. and Cortes, C. (2005) · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, E., Levy, K. Y., and Shalev-Shwartz, S. (2015) · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ji, Z. and Telgarsky, M. (2018) · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. (2018) · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. (2018) · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y. (2019) · 2019
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q. (2019) · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Characterizing the implicit bias via a primal-dual analysis
Ji, Z. and Telgarsky, M. (2019) · 2019
Earlier work this paper cites.
Convergence of gradient descent on separable data
Nacson, M. S., Lee, J., Gunasekar, S., Savarese, P. H. P., Srebro, N., and Soudry, D. (2019) · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. (2019) · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Earlier work this paper cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F. (2020) · 2020
Earlier work this paper cites.
Directional convergence and alignment in deep learning
Ji, Z. and Telgarsky, M. (2020) · 2020
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J. (2020) · 2020
Earlier work this paper cites.
Global convergence of deep networks with one wide layer followed by pyramidal topology
Nguyen, Q. N. and Mondelli, M. (2020) · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Cited alongside, same era.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Chatterji, N. S. and Long, P. M. (2021) · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A. (2021) · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Cited alongside, same era.
Proxy convexity: A unified framework for the analysis of neural networks trained by gradient descent
Frei, S. and Gu, Q. (2021) · 2021
Cited alongside, same era.
Openai: Introducing chatgpt
OpenAI (2022) · 2022
Later among the works it cites.
Unraveling attention via convex duality: Analysis and interpretations of vision transformers
Sahiner, A., Ergen, T., Ozturkler, B., Pauly, J., Mardani, M., and Pilanci, M. (2022) · 2022
Later among the works it cites.
On the implicit bias in deep-learning algorithms
Vardi, G. (2022) · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. (2022) · 2022
Later among the works it cites.
Binary classification of gaussian mixtures: Abundance of support vectors, benign overfitting, and regularization
Wang, K. and Thrampoulidis, C. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast margin maximization via dual acceleration
Ji, Z., Srebro, N., and Telgarsky, M. (2021) · 2021
Cited alongside, same era.
Characterizing the implicit bias via a primal-dual analysis
Ji, Z. and Telgarsky, M. (2021) · 2021
Cited alongside, same era.
Wilds: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., et al. (2021) · 2021
Cited alongside, same era.
Vision transformer for small-size datasets
Lee, S. H., Lee, S., and Song, B. C. (2021) · 2021
Cited alongside, same era.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Li, Z., Luo, Y., and Lyu, K. (2021) · 2021
Cited alongside, same era.
Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence
Loizou, N., Vaswani, S., Hadj Laradji, I., and Lacoste-Julien, S. (2021) · 2021
Cited alongside, same era.
The inductive bias of re{lu} networks on orthogonally separable data
Phuong, M. and Lampert, C. H. (2021) · 2021
Cited alongside, same era.
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. (2023) · 2023
Later among the works it cites.
Birth of a transformer: A memory viewpoint
Bietti, A., Cabannes, V., Bouchacourt, D., Jegou, H., and Bottou, L. (2023) · 2023
Later among the works it cites.
On the optimization and generalization of multi-head attention
Deora, P., Ghaderi, R., Taheri, H., and Thrampoulidis, C. (2023) · 2023
Later among the works it cites.
Benign overfitting in linear classifiers and leaky relu networks from kkt conditions for margin maximization
Frei, S., Vardi, G., Bartlett, P., and Srebro, N. (2023) · 2023
Later among the works it cites.
Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Kunstner, F., Chen, J., Lavington, J. W., and Schmidt, M. (2023) · 2023
Later among the works it cites.
Memorization capacity of multi-head attention in transformers
Mahdavi, S., Liao, R., and Thrampoulidis, C. (2023) · 2023
Later among the works it cites.
On the role of attention in prompt-tuning
Oymak, S., Rawat, A. S., Soltanolkotabi, M., and Thrampoulidis, C. (2023) · 2023
Later among the works it cites.
Representational strengths and limitations of transformers
Sanford, C., Hsu, D., and Telgarsky, M. (2023) · 2023
Later among the works it cites.
Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models
Xie, X., Zhou, P., Li, H., Lin, Z., and YAN, S. (2023) · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L. (2023) · 2023
Later among the works it cites.
From self-attention to markov models: Unveiling the dynamics of generative transformers
Ildiz, M. E., Huang, Y., Li, Y., Rawat, A. S., and Oymak, S. (2024) · 2024
Closest in time.
Mechanics of next token prediction with self-attention
Li, Y., Huang, Y., E Ildiz, M., Singh Rawat, A., and Oymak, S. (2024) · 2024
Closest in time.
Attention with markov: A framework for principled analysis of transformers via markov chains
Makkuva, A. V., Bondaschi, M., Girish, A., Nagle, A., Jaggi, M., Kim, H., and Gastpar, M. (2024) · 2024
Closest in time.
Momo: Momentum models for adaptive learning rates
Schaipp, F., Ohana, R., Eickenberg, M., Defazio, A., and Gower, R. M. (2024) · 2024
Closest in time.
Implicit regularization of gradient flow on one-layer softmax attention
Sheen, H., Chen, S., Wang, T., and Zhou, H. H. (2024) · 2024
Closest in time.
Implicit bias of next-token prediction
Thrampoulidis, C. (2024) · 2024
Closest in time.