Fetching the paper…
Reading the bibliography…
Transformer training is notoriously difficult, requiring a careful design of optimizers and use of various heuristics.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Earlier work this paper cites.
Lectures on convex optimization , volume 137
Yurii Nesterov et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding in: Proceedings of the 2019 conference of the north american chapter of the association for computational linguistics, 4171–4186.. acl
J Devlin, MW Chang, K Lee, and K Toutanova · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Earlier work this paper cites.
Robustness to unbounded smoothness of generalized signsgd
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang, and Zhenxun Zhuang · 2022
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Cited alongside, same era.
How does adaptive optimization impact local neural network geometry?
Kaiqi Jiang, Dhruv Malik, and Yuanzhi Li · 2022
Cited alongside, same era.
Unveiling transformers with lego: a synthetic reasoning task
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Cited alongside, same era.
A mechanism for sample-efficient in-context learning for sparse retrieval tasks
Jacob Abernethy, Alekh Agarwal, Teodor V Marinov, and Manfred K Warmuth · 2023
Cited alongside, same era.
The crucial role of normalization in sharpness-aware minimization
Yan Dai, Kwangjun Ahn, and Suvrit Sra · 2023
Closest in time.
In-context convergence of transformers
Yu Huang, Yuan Cheng, and Yingbin Liang · 2023
Closest in time.
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Closest in time.
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kwangjun Ahn, Sébastien Bubeck, Sinho Chewi, Yin Tat Lee, Felipe Suarez, and Yi Zhang · 2023
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Cited alongside, same era.
Physics of language models: Part 1, context-free grammar
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Hongkang Li, Meng Wang, Sijia Liu, and Pin-Yu Chen
Cited in the paper.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski
Cited in the paper.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie
Cited in the paper.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra
Cited in the paper.
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis · 2023
Closest in time.
Toward understanding why adam converges faster than sgd for transformers
Yan Pan and Yuanzhi Li · 2023
Closest in time.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
Closest in time.