Fetching the paper…
Reading the bibliography…
We initiate an investigation into the optimization properties of next-token prediction (NTP), the dominant training paradigm for modern language models.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
Prediction and entropy of printed english
Claude E Shannon · 1951
Earlier work this paper cites.
A note on one class of perceptrons
Vladimir N Vapnik and Alexey Ya Chervonenkis · 1964
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Thomas M Cover · 1965
Earlier work this paper cites.
Taking on the curse of dimensionality in joint distributions using neural networks
Samy Bengio and Yoshua Bengio · 2000
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Pattern classification
Peter E Hart, David G Stork, and Richard O Duda · 2000
Earlier work this paper cites.
Margin maximizing loss functions
Saharon Rosset, Ji Zhu, and Trevor J. Hastie · 2003
Earlier work this paper cites.
Lectures in geometric functional analysis
R. Vershynin · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Neural word embedding as implicit matrix factorization
Omer Levy and Yoav Goldberg · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Ambiguity helps: Classification with disagreements in crowdsourced annotations
Viktoriia Sharmanska, Daniel Hernández-Lobato, Jose Miguel Hernandez-Lobato, and Novi Quadrianto · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Does data interpolation contradict statistical optimality?
Mikhail Belkin, Alexander Rakhlin, and Alexandre B Tsybakov · 2018
Earlier work this paper cites.
Emmanuel J Candès and Pragya Sur · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Mor Shpigel Nacson, Nathan Srebro, and Daniel Soudry · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
A precise analysis of phasemax in phase retrieval
Fariborz Salehi, Ehsan Abbasi, and Babak Hassibi · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2019
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Earlier work this paper cites.
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2019
Earlier work this paper cites.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, and Daniel Soudry · 2019
Earlier work this paper cites.
Human uncertainty makes classification more robust
Joshua C Peterson, Ruairidh M Battleday, Thomas L Griffiths, and Olga Russakovsky · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
A modern maximum-likelihood theory for high-dimensional logistic regression
Pragya Sur and Emmanuel J Candès · 2019
Earlier work this paper cites.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Earlier work this paper cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Earlier work this paper cites.
Gradient descent follows the regularization path for general losses
Ziwei Ji, Miroslav Dudík, Robert E Schapire, and Matus Telgarsky · 2020
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Cited alongside, same era.
The role of regularization in classification of high-dimensional noisy gaussian mixture
Francesca Mignacco, Florent Krzakala, Yue M Lu, and Lenka Zdeborová · 2020
Cited alongside, same era.
Neural collapse with unconstrained features
Dustin G Mixon, Hans Parshall, and Jianzong Pi · 2020
Cited alongside, same era.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai · 2020
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, E. Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Later among the works it cites.
Matias D Cattaneo, Jason M Klusowski, and Boris Shigida · 2023
Later among the works it cites.
Benign overfitting in adversarially robust linear classification
Jinghui Chen, Yuan Cao, and Quanquan Gu · 2023
Later among the works it cites.
Learning curves for the multi-class teacher–student perceptron
Elisabetta Cornacchia, Francesca Mignacco, Rodrigo Veiga, Cédric Gerbelot, Bruno Loureiro, and Lenka Zdeborová · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An investigation of why overparameterization exacerbates spurious correlations
Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang · 2020
Cited alongside, same era.
On uniform convergence and low-norm interpolation learning
Lijia Zhou, Danica J Sutherland, and Nati Srebro · 2020
Cited alongside, same era.
Stochastic mirror descent on overparameterized nonlinear models
Navid Azizan, Sahin Lale, and Babak Hassibi · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Risk bounds for over-parameterized maximum margin classification on sub-gaussian mixtures
Yuan Cao, Quanquan Gu, and Mikhail Belkin · 2021
Cited alongside, same era.
When does gradient descent with logistic loss find interpolating two-layer networks?
Niladri S Chatterji, Philip M Long, and Peter L Bartlett · 2021
Cited alongside, same era.
Label noise sgd provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Cited alongside, same era.
On the optimization and generalization of multi-head attention
Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri, and Christos Thrampoulidis · 2023
Later among the works it cites.
A mathematical perspective on transformers
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Speech and Language Processing
Daniel Jurafsky and James H. Martin · 2023
Later among the works it cites.
Provable memorization capacity of transformers
Junghwan Kim, Michelle Kim, and Barzan Mozafari · 2023
Later among the works it cites.
Transformers as algorithms: Generalization and stability in in-context learning, 2023
Yingcong Li, M. Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Later among the works it cites.
Same pre-training loss, better downstream: Implicit bias matters for language models
Hong Liu, Sang Michael Xie, Zhiyuan Li, and Tengyu Ma · 2023
Later among the works it cites.
Auto-regressive next-token predictors are universal learners
Eran Malach · 2023
Later among the works it cites.
On the role of attention in prompt-tuning
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis · 2023
Later among the works it cites.
Asymptotic behavior of adversarial training in binary linear classification
Hossein Taheri, Ramtin Pedarsani, and Christos Thrampoulidis · 2023
Later among the works it cites.
Multinomial logistic regression: Asymptotic normality on null covariates in high-dimensions
Kai Tan and Pierre C Bellec · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Trained transformers learn linear models in-context, 2023
Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett · 2023
Later among the works it cites.
xlstm: Extended long short-term memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Closest in time.
The necessity of machine learning theory in mitigating ai risk
Mikhail Belkin · 2024
Closest in time.
Burak Çakmak, Yue M Lu, and Manfred Opper · 2024
Closest in time.
Provably learning a multi-head attention layer
Sitan Chen and Yuanzhi Li · 2024
Closest in time.
The evolution of statistical induction heads: In-context learning markov chains
Benjamin L Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
Closest in time.
Are transformers with one layer self-attention using low-rank weight matrices universal approximators?
Tokio Kajitsuka and Issei Sato · 2024
Closest in time.
Mechanics of next token prediction with self-attention
Yingcong Li, Yixiao Huang, Muhammed E Ildiz, Ankit Singh Rawat, and Samet Oymak · 2024
Closest in time.
Upper and lower memory capacity bounds of transformers for next-token prediction
Liam Madden, Curtis Fox, and Christos Thrampoulidis · 2024
Closest in time.
Attention with markov: A framework for principled analysis of transformers via markov chains
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Martin Jaggi, Hyeji Kim, and Michael Gastpar · 2024
Closest in time.
Implicit bias and fast convergence rates for self-attention
Bhavya Vasudeva, Puneesh Deora, and Christos Thrampoulidis · 2024
Closest in time.
Precise asymptotic generalization for multiclass classification with overparameterized linear models
David Wu and Anant Sahai · 2024
Closest in time.
Shuo Xie and Zhiyuan Li · 2024
Closest in time.
The implicit bias of adam on separable data
Chenyang Zhang, Difan Zou, and Yuan Cao · 2024
Closest in time.
Implicit geometry of next-token prediction: From language sparsity patterns to model representations
Yize Zhao, Tina Behnia, Vala Vakilian, and Christos Thrampoulidis · 2024
Closest in time.