Fetching the paper…
Reading the bibliography…
The attention mechanism is the key to large language models, and the attention matrix serves as an algorithmic and computational bottleneck for such a scheme.
Liii. on lines and planes of closest fit to systems of points in space
Karl Pearson · 1901
Earlier work this paper cites.
Linear regression analysis of economic time series
Tjalling Charles Koopmans · 1937
Earlier work this paper cites.
Least-squares fitting of a straight line
Derek York · 1966
Earlier work this paper cites.
An analysis of the total least squares problem
Gene H Golub and Charles F Van Loan · 1980
Earlier work this paper cites.
Estimation in a multivariate” errors in variables” regression model: large sample results
Leon Jay Gleser · 1981
Earlier work this paper cites.
Concepts for reliable modelling of linear systems with application to on-line identification of multivariable state space descriptions
J Staar · 1982
Earlier work this paper cites.
Extensions of lipschitz mappings into a hilbert space
William B Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
Analysis and solution of the nongeneric total least squares problem
Sabine Van Huffel and Joos Vandewalle · 1988
Earlier work this paper cites.
Faster approximation algorithms for the unit capacity concurrent flow problem with applications to routing and finding sparse cuts
Philip Klein, Serge Plotkin, Clifford Stein, and Eva Tardos · 1994
Earlier work this paper cites.
Latent semantic indexing: A probabilistic analysis
Christos H Papadimitriou, Hisao Tamaki, Prabhakar Raghavan, and Santosh Vempala · 1998
Earlier work this paper cites.
Authoritative sources in a hyperlinked environment
Jon M Kleinberg · 1999
Earlier work this paper cites.
Spectral analysis of data
Yossi Azar, Amos Fiat, Anna Karlin, Frank McSherry, and Jared Saia · 2001
Earlier work this paper cites.
Web search via hub synthesis
Dimitris Achlioptas, Amos Fiat, Anna R Karlin, and Frank McSherry · 2001
Earlier work this paper cites.
Spectral partitioning of random graphs
Frank McSherry · 2001
Earlier work this paper cites.
Competitive recommendation systems
Petros Drineas, Iordanis Kerenidis, and Prabhakar Raghavan · 2002
Earlier work this paper cites.
Clustering large graphs via the singular value decomposition
Petros Drineas, Alan Frieze, Ravi Kannan, Santosh Vempala, and Vishwanathan Vinay · 2004
Earlier work this paper cites.
On spectral learning of mixtures of distributions
Dimitris Achlioptas and Frank McSherry · 2005
Earlier work this paper cites.
Approximate nearest neighbors and the fast johnson-lindenstrauss transform
Nir Ailon and Bernard Chazelle · 2006
Earlier work this paper cites.
Improved approximation algorithms for large matrices via random projections
Tamas Sarlos · 2006
Earlier work this paper cites.
Fast approximation of centrality
Joseph Wang and David Eppstein · 2006
Earlier work this paper cites.
The spectral method for general mixture models
Ravindran Kannan, Hadi Salmasian, and Santosh Vempala · 2008
Earlier work this paper cites.
A fast randomized algorithm for the approximation of matrices
Franco Woolfe, Edo Liberty, Vladimir Rokhlin, and Mark Tygert · 2008
Earlier work this paper cites.
Blendenpik: Supercharging lapack’s least-squares solver
Haim Avron, Petar Maymounkov, and Sivan Toledo · 2010
Earlier work this paper cites.
Faster ridge regression via the subsampled randomized hadamard transform
Yichao Lu, Paramveer Dhillon, Dean P Foster, and Lyle Ungar · 2013
Earlier work this paper cites.
Low-distortion subspace embeddings in input-sparsity time and applications to robust linear regression
Xiangrui Meng and Michael W Mahoney · 2013
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Compressed matrix multiplication
Rasmus Pagh · 2013
Earlier work this paper cites.
Fast and scalable polynomial kernels via explicit feature maps
Ninh Pham and Rasmus Pagh · 2013
Earlier work this paper cites.
Subspace embeddings for the polynomial kernel
Haim Avron, Huy Nguyen, and David Woodruff · 2014
Earlier work this paper cites.
Near-optimal column-based matrix reconstruction
Christos Boutsidis, Petros Drineas, and Malik Magdon-Ismail · 2014
Earlier work this paper cites.
Sparser johnson-lindenstrauss transforms
Daniel M Kane and Jelani Nelson · 2014
Earlier work this paper cites.
Sketching as a tool for numerical linear algebra
David P Woodruff · 2014
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Large-scale random features for kernel regression
Valero Laparra, Diego Marcos Gonzalez, Devis Tuia, and Gustau Camps-Valls · 2015
Earlier work this paper cites.
Low-rank approximation and regression in input sparsity time
Kenneth L Clarkson and David P Woodruff · 2017
Earlier work this paper cites.
Fast Regression with an ℓ ∞ \ell_{\infty} Guarantee
Eric Price, Zhao Song, and David P. Woodruff · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Sketching for kronecker product regression and p-splines
Huaian Diao, Zhao Song, Wen Sun, and David Woodruff · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Cited alongside, same era.
Optimal sketching for kronecker product regression and low rank approximation
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Later among the works it cites.
A faster small treewidth sdp solver
Yuzhou Gu and Zhao Song · 2022
Later among the works it cites.
Attention mechanisms in computer vision: A survey
Meng-Hao Guo, Tian-Xing Xu, Jiang-Jiang Liu, Zheng-Ning Liu, Peng-Tao Jiang, Tai-Jiang Mu, Song-Hai Zhang, Ralph R Martin, Ming-Ming Cheng, and Shi-Min Hu · 2022
Later among the works it cites.
Capsule robot pose and mechanism state detection in ultrasound using attention-based hierarchical deep learning
Xiaoyun Liu, Daniel Esser, Brandon Wagstaff, Anna Zavodni, Naomi Matsuura, Jonathan Kelly, and Eric Diller · 2022
Later among the works it cites.
Adore: Differentially oblivious relational database operators
Lianke Qin, Rajesh Jayaram, Elaine Shi, Zhao Song, Danyang Zhuo, and Shumo Chu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huaian Diao, Rajesh Jayaram, Zhao Song, Wen Sun, and David Woodruff · 2019
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Cited alongside, same era.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Towards a zero-one law for column subset selection
Zhao Song, David Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Oblivious sketching of high-degree polynomial kernels
Thomas D Ahle, Michael Kapralov, Jakob BT Knudsen, Rasmus Pagh, Ameya Velingker, David P Woodruff, and Amir Zandieh · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Adaptive and dynamic multi-resolution hashing for pairwise summations
Lianke Qin, Aravind Reddy, Zhao Song, Zhaozhuo Xu, and Danyang Zhuo · 2022
Later among the works it cites.
Dynamic tensor product regression
Aravind Reddy, Zhao Song, and Lichen Zhang · 2022
Later among the works it cites.
pylspack: Parallel algorithms and data structures for sketching, column subset selection, regression, and leverage scores
Aleksandros Sobczyk and Efstratios Gallopoulos · 2022
Later among the works it cites.
Accelerating frank-wolfe algorithm using low-dimensional and adaptive data structures
Zhao Song, Zhaozhuo Xu, Yuanyuan Yang, and Lichen Zhang · 2022
Later among the works it cites.
Speeding up sparsification using inner product search data structures
Zhao Song, Zhaozhuo Xu, and Lichen Zhang · 2022
Later among the works it cites.
Faster algorithm for structured john ellipsoid computation
Zhao Song, Xin Yang, Yuanyuan Yang, and Tianyi Zhou · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2023
Closest in time.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Differentially private attention computation
Yeqi Gao, Zhao Song, and Xin Yang · 2023
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Fast quantum algorithm for attention computation
Yeqi Gao, Zhao Song, Xin Yang, and Ruizhe Zhang · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Space-efficient interior point method, with applications to linear programming and maximum weight bipartite matching
S. Cliff Liu, Zhao Song, Hengjie Zhang, Lichen Zhang, and Tianyi Zhou · 2023
Closest in time.
Fast submodular function maximization
Lianke Qin, Zhao Song, and Yitan Wang · 2023
Closest in time.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2023
Closest in time.
A general algorithm for solving rank-one matrix sensing
Lianke Qin, Zhao Song, and Ruizhe Zhang · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability
Zhao Song, Yitan Wang, Zheng Yu, and Lichen Zhang · 2023
Closest in time.
Sketching meets differential privacy: fast algorithm for dynamic kronecker projection maintenance
Zhao Song, Xin Yang, Yuanyuan Yang, and Lichen Zhang · 2023
Closest in time.
A nearly-optimal bound for fast regression with ℓ ∞ \ell_{\infty} guarantee
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2023
Closest in time.
A tale of two efficient value iteration algorithms for solving linear mdps with large action space
Zhaozhuo Xu, Zhao Song, and Anshumali Shrivastava · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.
Low rank matrix completion via robust alternating minimization in nearly linear time
Yuzhou Gu, Zhao Song, Junze Yin, and Lichen Zhang · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.