Fetching the paper…
Reading the bibliography…
In modern machine learning, attention computation is a fundamental task for training large language models such as Transformer, GPT-4 and ChatGPT.
Low rank approximation and regression in input sparsity time
Kenneth L. Clarkson and David P. Woodruff · 2013
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan Salakhutdinov, and Nathan Srebro · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P. Woodruff · 2016
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Low rank approximation with entrywise ℓ 1 \ell_{1} -norm error
Zhao Song, David P Woodruff, and Peilin Zhong · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Earlier work this paper cites.
Theoretical analysis of auto rate-tuning by batch normalization
Sanjeev Arora, Zhiyuan Li, and Kaifeng Lyu · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Sketching for kronecker product regression and p-splines
Huaian Diao, Zhao Song, Wen Sun, and David Woodruff · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size, 2018
Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Earlier work this paper cites.
A gram-gauss-newton method learning overparameterized deep neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Earlier work this paper cites.
Total least squares regression in input sparsity time
Huaian Diao, Zhao Song, David Woodruff, and Xin Yang · 2019
Earlier work this paper cites.
Norm matters: efficient and accurate normalization schemes in deep networks, 2019
Elad Hoffer, Ron Banner, Itay Golan, and Daniel Soudry · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Cited alongside, same era.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Cited alongside, same era.
Average case column subset selection for entrywise ℓ 1 \ell_{1} -norm loss
Zhao Song, David Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2021
Later among the works it cites.
Faster dynamic matrix inverse for faster lps
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2021
Later among the works it cites.
Oblivious sketching-based central path method for solving linear programming problems
Zhao Song and Zheng Yu · 2021
Later among the works it cites.
Approximating how single head attention learns
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Later among the works it cites.
Training multi-layer over-parametrized neural network in subquadratic time
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards a zero-one law for column subset selection
Zhao Song, David Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Relative error tensor low rank approximation
Zhao Song, David P Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Cited alongside, same era.
Solving tall dense linear programs in nearly linear time
Jan van den Brand, Yin Tat Lee, Aaron Sidford, and Zhao Song · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
A deterministic linear program solver in current matrix multiplication time
Jan van den Brand · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Later among the works it cites.
Colin Wei, Yining Chen, and Tengyu Ma · 2021
Later among the works it cites.
Toward understanding the feature learning process of self-supervised contrastive learning
Zixin Wen and Yuanzhi Li · 2021
Later among the works it cites.
Understanding the generalization of adam in learning neural networks with proper regularization
Difan Zou, Yuan Cao, Yuanzhi Li, and Quanquan Gu · 2021
Later among the works it cites.
Pixelated butterfly: Simple and efficient sparse training for neural network models
Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher Re · 2022
Later among the works it cites.
Towards understanding mixture of experts in deep learning
Zixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu, and Yuanzhi Li · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Discrepancy minimization in input-sparsity time
Yichuan Deng, Zhao Song, and Omri Weinstein · 2022
Later among the works it cites.
A faster small treewidth sdp solver
Yuzhou Gu and Zhao Song · 2022
Later among the works it cites.
Solving sdp faster: A robust ipm framework and efficient implementation
Baihe Huang, Shunhua Jiang, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Later among the works it cites.
Training overparametrized neural networks in sublinear time
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Later among the works it cites.
Towards understanding how momentum improves generalization in deep learning
Samy Jelassi and Yuanzhi Li · 2022
Later among the works it cites.
Adam is no better than normalized SGD: Dissecting how adaptivity improves GAN performance, 2022
Samy Jelassi, Arthur Mensch, Gauthier Gidel, and Yuanzhi Li · 2022
Later among the works it cites.
A faster interior-point method for sum-of-squares optimization, 2022
Shunhua Jiang, Bento Natura, and Omri Weinstein · 2022
Later among the works it cites.
Inductive bias of multi-channel linear convolutional networks with bounded weight norm
Meena Jagadeesan, Ilya Razenshteyn, and Suriya Gunasekar · 2022
Later among the works it cites.
Faster algorithm for structured john ellipsoid computation
Zhao Song, Xin Yang, Yuanyuan Yang, and Tianyi Zhou · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.