Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown their power in different areas.
The space complexity of approximating the frequency moments
Noga Alon, Yossi Matias, and Mario Szegedy · 1996
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Earlier work this paper cites.
Low-rank approximation and regression in input sparsity time
Kenneth L Clarkson and David P Woodruff · 2013
Earlier work this paper cites.
OSNAP: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Powers of tensors and fast matrix multiplication
François Le Gall · 2014
Earlier work this paper cites.
Unifying and strengthening hardness for dynamic problems via the online matrix-vector multiplication conjecture
Monika Henzinger, Sebastian Krinninger, Danupon Nanongkai, and Thatchaphol Saranurak · 2015
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P Woodruff · 2016
Earlier work this paper cites.
Faster online matrix-vector multiplication
Kasper Green Larsen and Ryan Williams · 2017
Earlier work this paper cites.
Low rank approximation with entrywise l1-norm error
Zhao Song, David P Woodruff, and Peilin Zhong · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Subspace embedding and linear regression with orlicz norm
Alexandr Andoni, Chengyu Lin, Ying Sheng, Peilin Zhong, and Ruiqi Zhong · 2018
Earlier work this paper cites.
Tight cell probe bounds for succinct boolean matrix-vector multiplication
Diptarka Chakraborty, Lior Kamma, and Kasper Green Larsen · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Dynamic matrix inverse: Improved algorithms and matching conditional lower bounds
Jan van den Brand, Danupon Nanongkai, and Thatchaphol Saranurak · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Earlier work this paper cites.
Hossein Esfandiari, Vahab Mirrokni, and Peilin Zhong · 2019
Earlier work this paper cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Earlier work this paper cites.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suarez, Yoann Dupont, Laurent Romary, Eric Villemonte de La Clergerie, Djame Seddah, and Benoit Sagot · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Relative error tensor low rank approximation
Zhao Song, David P Woodruff, and Peilin Zhong · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
A deterministic linear program solver in current matrix multiplication time
Jan van den Brand · 2020
Cited alongside, same era.
Smyrf-efficient attention using asymmetric clustering
Giannis Daras, Nikita Kitaev, Augustus Odena, and Alexandros G Dimakis · 2020
Cited alongside, same era.
An improved cutting plane method for convex optimization, convex-concave games and its applications
Haotian Jiang, Yin Tat Lee, Zhao Song, and Sam Chiu-wai Wong · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Massively parallel and dynamic algorithms for minimum size clustering
Alessandro Epasto, Mohammad Mahdian, Vahab Mirrokni, and Peilin Zhong · 2022
Later among the works it cites.
A faster small treewidth sdp solver
Yuzhou Gu and Zhao Song · 2022
Later among the works it cites.
Dynamic tensor product regression
Aravind Reddy, Zhao Song, and Lichen Zhang · 2022
Later among the works it cites.
Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability
Zhao Song, Yitan Wang, Zheng Yu, and Lichen Zhang · 2022
Later among the works it cites.
Accelerating frank-wolfe algorithm using low-dimensional and adaptive data structures
Zhao Song, Zhaozhuo Xu, Yuanyuan Yang, and Lichen Zhang · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Attention-based sentiment analysis using convolutional and recurrent neural network
Mohd Usama, Belal Ahmad, Enmin Song, M Shamim Hossain, Mubarak Alrashoud, and Ghulam Muhammad · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
Fast algorithm for solving structured convex programs
Guanghao Ye · 2020
Cited alongside, same era.
A refined laser method and faster matrix multiplication
Josh Alman and Virginia Vassilevska Williams · 2021
Cited alongside, same era.
Pixelated butterfly: Simple and efficient sparse training for neural network models
Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher Re · 2021
Cited alongside, same era.
Scatterbrain: Unifying sparse and low-rank attention
Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher Ré · 2021
Cited alongside, same era.
Later among the works it cites.
Speeding up sparsification using inner product search data structures
Zhao Song, Zhaozhuo Xu, and Lichen Zhang · 2022
Later among the works it cites.
Sketching meets differential privacy: Fast algorithm for dynamic kronecker projection maintenance
Zhao Song, Xin Yang, Yuanyuan Yang, and Lichen Zhang · 2022
Later among the works it cites.
Faster algorithm for structured john ellipsoid computation
Zhao Song, Xin Yang, Yuanyuan Yang, and Tianyi Zhou · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
An improved sample complexity for rank-1 matrix sensing
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Differentially private continual releases of streaming frequency moment estimations
Alessandro Epasto, Jieming Mao, Andres Munoz Medina, Vahab Mirrokni, Sergei Vassilvitskii, and Peilin Zhong · 2023
Closest in time.
Low rank matrix completion via robust alternating minimization in nearly linear time
Yuzhou Gu, Zhao Song, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Convex minimization with integer minima in O ~ ( n 4 ) \widetilde{O}(n^{4}) time
Haotian Jiang, Yin Tat Lee, Zhao Song, and Lichen Zhang · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
OpenAI · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
A nearly-optimal bound for fast regression with ℓ ∞ \ell_{\infty} guarantee
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.