Fetching the paper…
Reading the bibliography…
Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent.
Self-entrainment of a population of coupled non-linear oscillators
Yoshiki Kuramoto · 1975
Earlier work this paper cites.
Vlasov equations
Roland L’vovich Dobrushin · 1979
Earlier work this paper cites.
Mean shift, mode seeking, and clustering
Yizong Cheng · 1995
Earlier work this paper cites.
Novel type of phase transition in a system of self-driven particles
Tamás Vicsek, András Czirók, Eshel Ben-Jacob, Inon Cohen, and Ofer Shochet · 1995
Earlier work this paper cites.
K-plane clustering
Paul S Bradley and Olvi L Mangasarian · 2000
Earlier work this paper cites.
A discrete nonlinear and non-autonomous model of consensus
Ulrich Krause · 2000
Earlier work this paper cites.
Opinion dynamics and bounded confidence: models, analysis and simulation
Rainer Hegselmann and Ulrich Krause · 2002
Earlier work this paper cites.
The Kuramoto model: A simple paradigm for synchronization phenomena
Juan A Acebrón, Luis L Bonilla, Conrad J Pérez Vicente, Félix Ritort, and Renato Spigler · 2005
Earlier work this paper cites.
Emergent behavior in flocks
Felipe Cucker and Steve Smale · 2007
Earlier work this paper cites.
Subspace clustering
René Vidal · 2011
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
Mean field kinetic equations
François Golse · 2013
Earlier work this paper cites.
Transport equation with nonlocal velocity in Wasserstein spaces: convergence of numerical schemes
Benedetto Piccoli and Francesco Rossi · 2013
Earlier work this paper cites.
Clustering and asymptotic behavior in opinion formation
Pierre-Emmanuel Jabin and Sebastien Motsch · 2014
Cited alongside, same era.
Heterophilious dynamics enhances consensus
Sebastien Motsch and Eitan Tadmor · 2014
Cited alongside, same era.
Extremal laws for the real Ginibre ensemble
Brian Rider and Christopher D. Sinclair · 2014
Cited alongside, same era.
Control to flocking of the kinetic Cucker–Smale model
Benedetto Piccoli, Francesco Rossi, and Emmanuel Trélat · 2015
Cited alongside, same era.
Emergence of bi-cluster flocking for the Cucker–Smale model
Junghee Cho, Seung-Yeal Ha, Feimin Huang, Chunyin Jin, and Dongnam Ko · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Complete cluster predictability of the Cucker–Smale flocking model on the real line
Seung-Yeal Ha, Jeongho Kim, Jinyeong Park, and Xiongtao Zhang · 2019
Later among the works it cites.
A mean-field optimal control formulation of deep learning
E Weinan, Jiequn Han, and Qianxiao Li · 2019
Later among the works it cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Later among the works it cites.
Understanding and improving transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2020
Later among the works it cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, XY Han, and David L Donoho · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
On properties of the generalized Wasserstein distance
Benedetto Piccoli and Francesco Rossi · 2016
Cited alongside, same era.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A proposal on machine learning via dynamical systems
E Weinan · 2017
Cited alongside, same era.
Neural ordinary differential equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Cited alongside, same era.
Apoorv Vyas, Angelos Katharopoulos, and François Fleuret · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2020
Later among the works it cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2022
Later among the works it cites.
Sinkformers: Transformers with doubly stochastic attention
Michael E Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyré · 2022
Later among the works it cites.