Fetching the paper…
Reading the bibliography…
In the classical transformer attention scheme, we are given three $n \times d$ size matrices $Q, K, V$ (the query, key, and value tokens), and the goal is to compute a new $n \times d$ size matrix $D^{-1} \exp(QK^\top) V$ where $D = \mathrm{diag}( \exp(QK^\top) {\bf 1}_n )$.
Codes on algebraic curves
Valerii Denisovich Goppa · 1981
Earlier work this paper cites.
Modular curves, shimura curves, and goppa codes, better than varshamov-gilbert bound
Michael A Tsfasman, SG Vlădutx, and Th Zink · 1982
Earlier work this paper cites.
Trading group theory for randomness
László Babai · 1985
Earlier work this paper cites.
Private coins versus public coins in interactive proof systems
Shafi Goldwasser and Michael Sipser · 1986
Earlier work this paper cites.
A low-complexity construction of algebraic geometric codes better than the Gilbert-Varshamov bound
Kenneth Wing-Ki Shum · 2000
Earlier work this paper cites.
On the complexity of k-sat
Russell Impagliazzo and Ramamohan Paturi · 2001
Earlier work this paper cites.
A low-complexity algorithm for the construction of algebraic-geometric codes better than the gilbert-varshamov bound
Kenneth W Shum, Ilia Aleshnikov, P Vijay Kumar, Henning Stichtenoth, and Vinay Deolalikar · 2001
Earlier work this paper cites.
A new algorithm for optimal 2-constraint satisfaction and its implications
Ryan Williams · 2005
Earlier work this paper cites.
Computational complexity: a modern approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
Algebrization: A new barrier in complexity theory
Scott Aaronson and Avi Wigderson · 2009
Earlier work this paper cites.
Lecture notes 8 of madhu sudan’s class, and scribed by josh alman
Madhu Sudan · 2013
Earlier work this paper cites.
Consequences of faster alignment of sequences
Amir Abboud, Virginia Vassilevska Williams, and Oren Weimann · 2014
Earlier work this paper cites.
Smoothed analysis of tensor decompositions
Aditya Bhaskara, Moses Charikar, Ankur Moitra, and Aravindan Vijayaraghavan · 2014
Earlier work this paper cites.
Distributed pcp theorems for hardness of approximation in p
Amir Abboud, Aviad Rubinstein, and Ryan Williams · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Efficient density evaluation for smooth kernels
Arturs Backurs, Moses Charikar, Piotr Indyk, and Paris Siminelakis · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Sketching for kronecker product regression and p-splines
Huaian Diao, Zhao Song, Wen Sun, and David Woodruff · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Hardness of approximate nearest neighbor search
Aviad Rubinstein · 2018
Earlier work this paper cites.
On some fine-grained questions in algorithms and complexity
Virginia Vassilevska Williams · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Optimal sketching for kronecker product regression and low rank approximation
Huaian Diao, Rajesh Jayaram, Zhao Song, Wen Sun, and David Woodruff · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Viji Srinivasan, Xiaodong Cui, Wei Zhang, and Kailash Gopalakrishnan · 2019
Earlier work this paper cites.
Relative error tensor low rank approximation
Zhao Song, David P Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Greg Yang · 2019
Cited alongside, same era.
Q8bert: Quantized 8bit bert
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat · 2019
Cited alongside, same era.
Algorithms and hardness for linear algebra on geometric graphs
Josh Alman, Timothy Chu, Aaron Schild, and Zhao Song · 2020
Cited alongside, same era.
Oblivious sketching of high-degree polynomial kernels
Thomas D Ahle, Michael Kapralov, Jakob BT Knudsen, Rasmus Pagh, Ameya Velingker, David P Woodruff, and Amir Zandieh · 2020
Cited alongside, same era.
Smoothed analysis for tensor methods in unsupervised learning
Aditya Bhaskara, Aidao Chen, Aidan Perreault, and Aravindan Vijayaraghavan · 2020
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Closest in time.
Faster robust tensor power method for arbitrary order
Yichuan Deng, Zhao Song, and Junze Yin · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Smyrf-efficient attention using asymmetric clustering
Giannis Daras, Nikita Kitaev, Augustus Odena, and Alexandros G Dimakis · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2020
Cited alongside, same era.
Closest in time.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Closest in time.
Differentially private attention computation
Yeqi Gao, Zhao Song, and Xin Yang · 2023
Closest in time.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Praneeth Kacham, Vahab Mirrokni, and Peilin Zhong · 2023
Closest in time.
On the computational complexity of self-attention
Feyza Duman Keles, Pruthuvi Mahesakya Wijewardena, and Chinmay Hegde · 2023
Closest in time.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhuang Yuan, Zhao Song, Anshumali Shrivastava, Yuandong Tian Ce Zhang, Christopher Re, and Beidi Chen · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Training and inference of large language models using 8-bit floating point
Sergio P. Perez, Yan Zhang, James Briggs, Charlie Blake, Josh Levy-Kramer, Paul Balanca, Carlo Lushi, Stephen Barlow, and Andrew Fitzgibbon · 2023
Closest in time.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Closest in time.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Closest in time.
Efficient post-training quantization with fp8 formats
Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, and Mengni Wang · 2023
Closest in time.
Sketching meets differential privacy: fast algorithm for dynamic kronecker projection maintenance
Zhao Song, Xin Yang, Yuanyuan Yang, and Lichen Zhang · 2023
Closest in time.
A nearly-optimal bound for fast regression with ℓ ∞ \ell_{\infty} guarantee
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Streaming semidefinite programs: O( n \sqrt{n} ) passes, small space and fast runtime
Zhao Song, Mingquan Ye, and Lichen Zhang · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.
H 2 O : Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen · 2023
Closest in time.