Fetching the paper…
Reading the bibliography…
Since its inception in "Attention Is All You Need", transformer architecture has led to revolutionary advancements in NLP.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Atomic decomposition by basis pursuit
Scott Shaobing Chen, David L Donoho, and Michael A Saunders · 2001
Earlier work this paper cites.
Matrix rank minimization with applications
Maryam Fazel · 2002
Earlier work this paper cites.
Margin maximizing loss functions
Saharon Rosset, Ji Zhu, and Trevor Hastie · 2003
Earlier work this paper cites.
Maximum-margin matrix factorization
Nathan Srebro, Jason Rennie, and Tommi Jaakkola · 2004
Earlier work this paper cites.
Boosting with early stopping: Convergence and consistency
Tong Zhang and Bin Yu · 2005
Earlier work this paper cites.
Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information
Emmanuel J Candès, Justin Romberg, and Terence Tao · 2006
Earlier work this paper cites.
Compressed sensing
David L Donoho · 2006
Earlier work this paper cites.
Signal recovery from random measurements via orthogonal matching pursuit
Joel A Tropp and Anna C Gilbert · 2007
Earlier work this paper cites.
New null space results and recovery thresholds for matrix rank minimization
Samet Oymak and Babak Hassibi · 2010
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
A simplified approach to recovery conditions for low rank matrices
Samet Oymak, Karthik Mohan, Maryam Fazel, and Babak Hassibi · 2011
Earlier work this paper cites.
Null space conditions and thresholds for rank minimization
Benjamin Recht, Weiyu Xu, and Babak Hassibi · 2011
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
Low-rank solutions of linear matrix equations via procrustes flow
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Ben Recht · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Connecting optimization and regularization paths
Arun Suggala, Adarsh Prasad, and Pradeep K Ravikumar · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Earlier work this paper cites.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Earlier work this paper cites.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, and Daniel Soudry · 2019
Cited alongside, same era.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Cited alongside, same era.
The implicit bias of adagrad on separable data
Qian Qian and Xiaoyuan Qian · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
Winnowing with gradient descent
Ehsan Amid and Manfred K Warmuth · 2020
Cited alongside, same era.
Reparameterizing mirror descent as gradient descent
Ehsan Amid and Manfred KK Warmuth · 2020
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
The implicit bias for adaptive optimization algorithms on homogeneous neural networks
Bohan Wang, Qi Meng, Wei Chen, and Tie-Yan Liu · 2021
Later among the works it cites.
Momentum doesn’t change the implicit bias
Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun, Wei Chen, and Zhi-Ming Ma · 2021
Later among the works it cites.
The benefits of implicit regularization from sgd in least squares problems
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, Dean P Foster, and Sham Kakade · 2021
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Cited alongside, same era.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and et al · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason D Lee, and Tengyu Ma · 2020
Cited alongside, same era.
Gradient descent follows the regularization path for general losses
Ziwei Ji, Miroslav Dudík, Robert E Schapire, and Matus Telgarsky · 2020
Cited alongside, same era.
Later among the works it cites.
Pierre Baldi and Roman Vershynin · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Later among the works it cites.
Convexifying transformers: Improving optimization and understanding of transformer networks
Tolga Ergen, Behnam Neyshabur, and Harsh Mehta · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Eli Sander, and Yuanzhi Li · 2022
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
OpenAI: Introducing ChatGPT
OpenAI · 2022
Later among the works it cites.
Mirror descent maximizes generalized margin and can be implemented efficiently
Haoyuan Sun, Kwangjun Ahn, Christos Thrampoulidis, and Navid Azizan · 2022
Later among the works it cites.
Unraveling attention via convex duality: Analysis and interpretations of vision transformers
Arda Sahiner, Tolga Ergen, Batu Ozturkler, John Pauly, Morteza Mardani, and Mert Pilanci · 2022
Later among the works it cites.
Imbalance trouble: Revisiting neural-collapse geometry
Christos Thrampoulidis, Ganesh Ramachandra Kini, Vala Vakilian, and Tina Behnia · 2022
Later among the works it cites.
Binary classification of gaussian mixtures: Abundance of support vectors, benign overfitting, and regularization
Ke Wang and Christos Thrampoulidis · 2022
Later among the works it cites.
Flowformer: Linearizing transformers with conservation flows
Haixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Closest in time.
Transformers learn through gradual rank increase
Enric Boix-Adsera, Etai Littwin, Emmanuel Abbe, Samy Bengio, and Joshua Susskind · 2023
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Closest in time.
Jigsaw-vit: Learning jigsaw puzzles in vision transformer
Yingyi Chen, Xi Shen, Yahui Liu, Qinghua Tao, and Johan AK Suykens · 2023
Closest in time.
What can a single attention layer learn? a study through the random features lens
Hengyu Fu, Tianyu Guo, Yu Bai, and Song Mei · 2023
Closest in time.
Try bard, an ai experiment by google
Google · 2023
Closest in time.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, M Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Closest in time.
Hongkang Li, Meng Wang, Sijia Liu, and Pin-Yu Chen · 2023
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He, Thomas Hofmann, Chris Maddison, and Daniel M Roy · 2023
Closest in time.
A primal-dual framework for transformers and neural networks
Tan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L Bertozzi, Richard Baraniuk, and Stanley Osher · 2023
Closest in time.
OpenAI · 2023
Closest in time.
On the role of attention in prompt-tuning
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Margin maximization in attention mechanism
Davoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, and Samet Oymak · 2023
Closest in time.
Implicit regularization towards rank minimization in relu networks
Nadav Timor, Gal Vardi, and Ohad Shamir · 2023
Closest in time.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Du · 2023
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
Closest in time.