Fetching the paper…
Reading the bibliography…
Given a sequence of tokens, such as words, the task of next-token prediction is to predict the next-token conditional probability distribution.
M. Fréchet, “Des familles et fonctions additivies d’ensembles abstraits,” Fundamenta Mathematicae , vol. 5, pp. 206–251, 1924
1924
Earlier work this paper cites.
A. Sard, “The measure of the critical values of differentiable maps,” Bulletin of the American Mathematical Society , vol. 48, pp. 883–890, 1942
1942
Earlier work this paper cites.
C. E. Shannon, “A mathematical theory of communication,” The Bell System Technical Journal , vol. 27, pp. 379–423, 623–656, 1948
1948
Earlier work this paper cites.
——, “Prediction and entropy of printed english,” The Bell System Technical Journal , vol. 30, no. 1, pp. 50–64, 1951
1951
Earlier work this paper cites.
F. R. Gantmacher, The Theory of Matrices, Volume 1 . New York: Chelsea Publishing Company, 1960
1960
Earlier work this paper cites.
T. M. Cover, “Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition,” IEEE transactions on electronic computers , vol. 3, pp. 326–334, 1965
1965
Earlier work this paper cites.
R. C. Gunning and H. Rossi, Analytic functions of several complex variables . Englewood Cliffs, N.J.: Prentice-Hall, Inc., 1965
1965
Earlier work this paper cites.
T. L. Booth and R. A. Thompson, “Applying probability measures to abstract languages,” IEEE Transactions on Computers , vol. C-22, pp. 442–450, 1973
1973
Earlier work this paper cites.
J. B. Kruskal, “Three-way arrays: rank and uniqueness of trilinear decompositions, with application to arithmetic complexity and statistics,” Linear Algebra and its Applications , vol. 18, no. 2, pp. 95–138, 1977
1977
Earlier work this paper cites.
E. B. Baum, “On the capabilities of multilayer perceptrons,” Journal of Complexity , vol. 4, no. 3, pp. 193–215, 1988
1988
Earlier work this paper cites.
S.-C. Huang and Y.-F. Huang, “Bounds on the number of hidden neurons in multilayer perceptrons,” IEEE Transactions on Neural Networks , vol. 2, no. 1, pp. 47–55, 1991
1991
Earlier work this paper cites.
S. Tamura, “Capabilities of a three layer feedforward neural network,” IEEE International Joint Conference on Neural Networks , vol. 3, pp. 2757–2762, 1991
1991
Earlier work this paper cites.
A. Sakurai, “n-h-1 networks store no less n*h+1 examples, but sometimes no more,” in International Joint Conference on Neural Networks (IJCNN) , vol. 3, 1992, pp. 936–941
1992
Earlier work this paper cites.
M. Yamasaki, “The lower bound of the capacity for a neural network with multiple hidden layers,” in International Conference on Artificial Neural Networks (ICANN) , 1993, pp. 546–549
1993
Earlier work this paper cites.
L. van den Dries, A. Macintyre, and D. Marker, “The elementary theory of restricted analytic fields with exponentiation,” Annals of Mathematics , vol. 140, pp. 183–205, 1994
1994
Earlier work this paper cites.
L. van den Dries and C. Miller, “On the real exponential field with restricted analytic functions,” Israel Journal of Mathematics , vol. 85, pp. 19–56, 1994
1994
Earlier work this paper cites.
P. Gage, “A new algorithm for data compression,” The C Users Journal , vol. 12, no. 2, p. 23–38, 1994
1994
Earlier work this paper cites.
P. Billingsley, Probability and Measure , 3rd ed. John Wiley & Sons, 1995
1995
Earlier work this paper cites.
E. D. Sontag, “Critical points for least-squares problems involving certain analytic functions, with applications to sigmoidal nets,” Advances in Computational Mathematics , vol. 5, no. 1, pp. 245–268, 1996
1996
Earlier work this paper cites.
E. D. Sontag, “Shattering all sets of k points in “general position” requires (k - 1)/2 parameters,” Neural Computation , vol. 9, no. 2, pp. 337–348, 1997
1997
Cited alongside, same era.
S. Tamura and M. Tateishi, “Capabilities of a four-layered feedforward neural network: four layers versus three,” IEEE Transactions on Neural Networks , vol. 8, no. 2, pp. 251–255, 1997
1997
Cited alongside, same era.
G.-B. Huang and H. Babri, “Upper bounds on the number of hidden neurons in feedforward networks with arbitrary bounded nonlinear activation functions,” IEEE Transactions on Neural Networks , vol. 9, no. 1, pp. 224–229, 1998
1998
Cited alongside, same era.
L. van den Dries, Tame Topology and O-minimal Structures . Cambridge University Press, 1998
1998
Cited alongside, same era.
Y. Bengio, R. Ducharme, and P. Vincent, “A neural probabilistic language model,” in Neural Information Processing Systems (NeurIPS) , vol. 13, 2000
Y. Tian, Y. Wang, B. Chen, and S. S. Du, “Scan and snap: Understanding training dynamics and token composition in 1-layer transformer,” in Neural Information Processing Systems (NeurIPS) , 2023
2023
Later among the works it cites.
D. A. Tarzanagh, Y. Li, C. Thrampoulidis, and S. Oymak, “Transformers as support vector machines,” in NeurIPS 2023 Workshop on Mathematics of Modern Machine Learning , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Sanford, D. J. Hsu, and M. Telgarsky, “Representational strengths and limitations of transformers,” in Neural Information Processing Systems (NeurIPS) , vol. 36, 2023, pp. 36 677–36 707
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2000
Cited alongside, same era.
G.-B. Huang, “Learning capability and storage capacity of two-hidden-layer feedforward networks,” IEEE Transactions on Neural Networks , vol. 14, no. 2, pp. 274–281, 2003
2003
Cited alongside, same era.
A. A. Markov, “An example of statistical investigation of the text eugene onegin concerning the connection of samples in chains,” Science in Context , vol. 19, no. 4, p. 591–600, 2006
2006
Cited alongside, same era.
T. Mikolov, “Statistical language models based on neural networks,” Ph.D. thesis, Brno University of Technology, Faculty of Information Technology, 2012
2012
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning: From Theory to Algorithms . Cambridge University Press, 2014
2014
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in International Conference on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
S. Axler, Linear Algebra Done Right, Third Edition . New York: Springer, 2015
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
2023
Later among the works it cites.
L. Madden and C. Thrampoulidis, “Memory capacity of two layer neural networks with smooth activations,” SIAM Journal on Mathematics of Data Science , vol. 6, no. 3, pp. 679–702, 2024
2024
Closest in time.
2024
Closest in time.
T. Kajitsuka and I. Sato, “Are transformers with one layer self-attention using low-rank weight matrices universal approximators?” in Conference on Learning Theory (COLT) , 2024
2024
Closest in time.
S. Mahdavi, R. Liao, and C. Thrampoulidis, “Memorization capacity of multi-head attention in transformers,” in International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
2024
Closest in time.
R. Zhang, S. Frei, and P. L. Bartlett, “Trained transformers learn linear models in-context,” Journal of Machine Learning Research (JMLR) , vol. 25, no. 49, pp. 1–55, 2024
2024
Closest in time.
P. Deora, R. Ghaderi, H. Taheri, and C. Thrampoulidis, “On the optimization and generalization of multi-head attention,” Transactions on Machine Learning Research (TMLR) , 2024
2024
Closest in time.
2024
Closest in time.
Y. Li, Y. Huang, M. E. Ildiz, A. S. Rawat, and S. Oymak, “Mechanics of next token prediction with self-attention,” in International Conference on Artificial Intelligence and Statistics (AISTATS) , 2024
2024
Closest in time.
Y. Tian, Y. Wang, Z. Zhang, B. Chen, and S. S. Du, “JoMA: Demystifying multilayer transformers via joint dynamics of MLP and attention,” in International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
2024
Closest in time.
C. Thrampoulidis, “Implicit bias of next-token prediction,” arXiv preprint arXiv:2402.18551 , 2024
2024
Closest in time.
Z. Wang, S. Wei, D. Hsu, and J. D. Lee, “Transformers provably learn sparse token selection while fully-connected nets cannot,” in International Conference on Machine Learning (ICML) , vol. 235, 2024, pp. 51 854–51 912
2024
Closest in time.