Fetching the paper…
Reading the bibliography…
Transformers play a central role in the inner workings of large language models.
Orthogonal polynomials
Gabor Szegö · 1939
Earlier work this paper cites.
A problem in geometric probability
James G Wendel · 1962
Earlier work this paper cites.
Une propriété topologique des sous-ensembles analytiques réels
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
The Speed of Mean Glivenko-Cantelli Convergence
R. M. Dudley · 1969
Earlier work this paper cites.
Self-entrainment of a population of coupled non-linear oscillators
Yoshiki Kuramoto · 1975
Earlier work this paper cites.
Vlasov equations
Roland L’vovich Dobrushin · 1979
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Spherical codes and designs
Philippe Delsarte, Jean-Marie Goethals, and Johan Jacob Seidel · 1991
Earlier work this paper cites.
Order function and macroscopic mutual entrainment in uniformly coupled limit-cycle oscillators
Hiroaki Daido · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Novel type of phase transition in a system of self-driven particles
Tamás Vicsek, András Czirók, Eshel Ben-Jacob, Inon Cohen, and Ofer Shochet · 1995
Earlier work this paper cites.
The variational formulation of the Fokker–Planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
A computational fluid mechanics solution to the Monge-Kantorovich mass transfer problem
Jean-David Benamou and Yann Brenier · 2000
Earlier work this paper cites.
A discrete nonlinear and non-autonomous model of consensus
Ulrich Krause · 2000
Earlier work this paper cites.
From Kuramoto to Crawford: exploring the onset of synchronization in populations of coupled oscillators
Steven H Strogatz · 2000
Earlier work this paper cites.
The geometry of dissipative evolution equations: the porous medium equation
Felix Otto · 2001
Earlier work this paper cites.
Limite de champ moyen
Cédric Villani · 2001
Earlier work this paper cites.
Opinion dynamics and bounded confidence: models, analysis and simulation
Rainer Hegselmann and Ulrich Krause · 2002
Earlier work this paper cites.
Upper bounds on coarsening rates
Robert V Kohn and Felix Otto · 2002
Earlier work this paper cites.
The Kuramoto model: A simple paradigm for synchronization phenomena
Juan A Acebrón, Luis L Bonilla, Conrad J Pérez Vicente, Félix Ritort, and Renato Spigler · 2005
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré · 2005
Earlier work this paper cites.
Universally optimal distribution of points on spheres
Henry Cohn and Abhinav Kumar · 2007
Earlier work this paper cites.
Emergent behavior in flocks
Felipe Cucker and Steve Smale · 2007
Earlier work this paper cites.
Slow motion of gradient flows
Felix Otto and Maria G Reznikoff · 2007
Earlier work this paper cites.
Infinite time aggregation for the critical Patlak-Keller-Segel model in ℝ 2 \mathbb{R}^{2}
Adrien Blanchet, José A Carrillo, and Nader Masmoudi · 2008
Earlier work this paper cites.
Experimental study of energy-minimizing point configurations on spheres
Brandon Ballinger, Grigoriy Blekherman, Henry Cohn, Noah Giansiracusa, Elizabeth Kelly, and Achill Schürmann · 2009
Earlier work this paper cites.
Optimal transport: old and new
Cédric Villani · 2009
Earlier work this paper cites.
L p L^{p} theory for the multidimensional aggregation equation
Andrea L Bertozzi, Thomas Laurent, and Jesús Rosado · 2011
Earlier work this paper cites.
Global-in-time weak measure solutions and finite-time aggregation for nonlocal interaction equations
J. A. Carrillo, M. DiFrancesco, A. Figalli, T. Laurent, and D. Slepčev · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
On the empirical estimation of integral probability metrics
Bharath K. Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Gert R. G. Lanckriet · 2012
Earlier work this paper cites.
There is no non-zero stable fixed point for dense networks in the homogeneous Kuramoto model
Richard Taylor · 2012
Earlier work this paper cites.
Approximation theory and harmonic analysis on spheres and balls
Feng Dai and Yuan Xu · 2013
Earlier work this paper cites.
From Newton to Boltzmann: hard spheres and short-range potentials
Isabelle Gallagher, Laure Saint-Raymond, and Benjamin Texier · 2013
Earlier work this paper cites.
Global stability of dynamical systems
Michael Shub · 2013
Earlier work this paper cites.
On the mean speed of convergence of empirical and occupation measures in Wasserstein distance
Emmanuel Boissard and Thibaut Le Gouic · 2014
Earlier work this paper cites.
Contractivity of transport distances for the kinetic Kuramoto equation
José A Carrillo, Young-Pil Choi, Seung-Yeal Ha, Moon-Jin Kang, and Yongduck Kim · 2014
Earlier work this paper cites.
Clustering and asymptotic behavior in opinion formation
Pierre-Emmanuel Jabin and Sebastien Motsch · 2014
Earlier work this paper cites.
Heterophilious dynamics enhances consensus
Sebastien Motsch and Eitan Tadmor · 2014
Earlier work this paper cites.
On the complete phase synchronization for the Kuramoto model in the mean-field limit
Dario Benedetto, Emanuele Caglioti, and Umberto Montemagno · 2015
Earlier work this paper cites.
A proof of the Kuramoto conjecture for a bifurcation structure of the infinite-dimensional Kuramoto model
Hayato Chiba · 2015
Earlier work this paper cites.
A nonlinear model of opinion formation on the sphere
Marco Caponigro, Anna Chiara Lai, and Benedetto Piccoli · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Landau damping in the Kuramoto model
Bastien Fernandez, David Gérard-Varet, and Giambattista Giacomin · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
On the dynamics of large particle systems in the mean field limit
François Golse · 2016
Cited alongside, same era.
Collective synchronization of classical and quantum oscillators
Seung-Yeal Ha, Dongnam Ko, Jinyeong Park, and Xiongtao Zhang · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
A proposal on machine learning via dynamical systems
Weinan E · 2017
Cited alongside, same era.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Cited alongside, same era.
Deep learning via dynamical systems: An approximation perspective
Qianxiao Li, Ting Lin, and Zuowei Shen · 2022
Later among the works it cites.
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2022
Later among the works it cites.
On the trend to global equilibrium for Kuramoto oscillators
Javier Morales and David Poyato · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Later among the works it cites.
Formal algorithms for transformers
Mary Phuong and Marcus Hutter · 2022
Later among the works it cites.
Sinkformers: Transformers with doubly stochastic attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Johan Markdahl, Johan Thunberg, and Jorge Gonçalves · 2017
Cited alongside, same era.
A consensus-based model for global optimization and its mean-field limit
René Pinnau, Claudia Totzeck, Oliver Tse, and Stephan Martin · 2017
Cited alongside, same era.
Energy optimization for distributions on the sphere and improvement to the Welch bounds
Yan Shuo Tan · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Neural ordinary differential equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Cited alongside, same era.
Landau damping to partially locked states in the Kuramoto model
Helge Dietert, Bastien Fernandez, and David Gérard-Varet · 2018
Cited alongside, same era.
Michael E Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyré · 2022
Later among the works it cites.
Universal approximation power of deep residual neural networks through the lens of control
Paulo Tabuada and Bahman Gharesifard · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Closest in time.
Sumformer: Universal Approximation for Efficient Transformers
Silas Alberti, Niclas Dern, Laura Thesing, and Gitta Kutyniok · 2023
Closest in time.
Interpolation, approximation and controllability of deep neural networks
Jingpu Cheng, Qianxiao Li, Ting Lin, and Zuowei Shen · 2023
Closest in time.
Sparsity in long-time control of neural ODEs
Carlos Esteve-Yagüe and Borjan Geshkovski · 2023
Closest in time.
Contranorm: A contrastive learning perspective on oversmoothing and beyond
Xiaojun Guo, Yifei Wang, Tianqi Du, and Yisen Wang · 2023
Closest in time.
A class of dimension-free metrics for the convergence of empirical measures
Jiequn Han, Ruimeng Hu, and Jihao Long · 2023
Closest in time.
On the impact of activation and normalization in obtaining isometric embeddings at initialization
Amir Joudaki, Hadi Daneshmand, and Francis Bach · 2023
Closest in time.
Approximation theory of transformer networks for sequence modeling
Haotian Jiang and Qianxiao Li · 2023
Closest in time.
A brief survey on the approximation theory for sequence modelling
Haotian Jiang, Qianxiao Li, Zhong Li, and Shida Wang · 2023
Closest in time.
Hierarchies, entropy, and quantitative propagation of chaos for mean field diffusions
Daniel Lacker · 2023
Closest in time.
Sharp uniform-in-time propagation of chaos
Daniel Lacker and Luc Le Flem · 2023
Closest in time.
Neural ODE control for classification, approximation, and transport
Domenec Ruiz-Balet and Enrique Zuazua · 2023
Closest in time.
Global-in-time mean-field convergence for singular Riesz-type diffusive flows
Matthew Rosenzweig and Sylvia Serfaty · 2023
Closest in time.
Token contrast for weakly-supervised semantic segmentation
Lixiang Ru, Heliang Zheng, Yibing Zhan, and Bo Du · 2023
Closest in time.
Swarming: hydrodynamic alignment with pressure
Eitan Tadmor · 2023
Closest in time.
Transformers as support vector machines
Davoud Ataee Tarzanagh, Yingcong Li, Christos Thrampoulidis, and Samet Oymak · 2023
Closest in time.
Nonlinear controllability and function representation by neural stochastic differential equations
Tanya Veeravalli and Maxim Raginsky · 2023
Closest in time.
Introduction to Transformers: an NLP Perspective
Tong Xiao and Jingbo Zhu · 2023
Closest in time.
Stabilizing transformer training by preventing attention entropy collapse
Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang, Jiatao Gu, and Joshua M Susskind · 2023
Closest in time.
Are more layers beneficial to graph transformers?
Haiteng Zhao, Shuming Ma, Dongdong Zhang, Zhi-Hong Deng, and Furu Wei · 2023
Closest in time.
Approximate controllability of continuity equation of transformers
Daniel Owusu Adu and Bahman Gharesifard · 2024
Closest in time.
Andrei Agrachev and Cyril Letrouit · 2024
Closest in time.
Self-attention networks localize when qk-eigenspectrum concentrates
Han Bao, Ryuichiro Hataya, and Ryo Karakida · 2024
Closest in time.
Geometric dynamics of signal propagation predict trainability of transformers
Aditya Cowsik, Tamra Nebabu, Xiao-Liang Qi, and Surya Ganguli · 2024
Closest in time.
Sinho Chewi, Jonathan Niles-Weed, and Philippe Rigollet · 2024
Closest in time.
Synchronization on circles and spheres with nonlinear interactions, 2024
Christopher Criscitiello, Quentin Rebjock, Andrew D. McRae, and Nicolas Boumal · 2024
Closest in time.
Setting the record straight on transformer oversmoothing
Gbètondji JS Dovonon, Michael M Bronstein, and Matt J Kusner · 2024
Closest in time.
On the optimization and generalization of multi-head attention
Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri, and Christos Thrampoulidis · 2024
Closest in time.
Transformers are universal in-context learners
Takashi Furuya, Maarten V de Hoop, and Gabriel Peyré · 2024
Closest in time.
Uniform in time propagation of chaos for the 2D vortex model and other singular stochastic systems
Arnaud Guillin, Pierre Le Bris, and Pierre Monmarché · 2024
Closest in time.
The emergence of clusters in self-attention dynamics
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet · 2024
Closest in time.
https://github.com/mistralai/mistral-finetune/blob/main/model/transformer.py , 2024
MistralAI · 2024
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Li, Bobby He, Thomas Hofmann, Chris J Maddison, and Dan Roy · 2024
Closest in time.
https://github.com/openai/gpt-2/blob/master/src/model.py , 2024
OpenAI · 2024
Closest in time.
Residual connections and normalization can provably prevent oversmoothing in gnns
Michael Scholkemper, Xinyi Wu, Ali Jadbabaie, and Michael Schaub · 2024
Closest in time.
On the role of attention masks and layernorm in transformers
Xinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka, and Ali Jadbabaie · 2024
Closest in time.
Demystifying oversmoothing in attention-based graph neural networks
Xinyi Wu, Amir Ajorlou, Zihui Wu, and Ali Jadbabaie · 2024
Closest in time.