Fetching the paper…
Reading the bibliography…
Sparse Attention is a technique that approximates standard attention computation with sub-quadratic complexity.
Über dyadische brüche
Aleksandr Khintchine · 1923
Earlier work this paper cites.
On a modification of chebyshev’s inequality and of the error formula of laplace
Sergei Bernstein · 1924
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff · 1952
Earlier work this paper cites.
A bound on tail probabilities for quadratic forms in independent random variables
David Lee Hanson and Farroll Tim Wright · 1971
Earlier work this paper cites.
The best constants in the khintchine inequality
Uffe Haagerup · 1981
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1994
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
An introduction to heavy-tailed and subexponential distributions
Sergey Foss, Dmitry Korshunov, Stan Zachary, et al · 2011
Earlier work this paper cites.
Improved analysis of the subsampled randomized hadamard transform
Joel A Tropp · 2011
Earlier work this paper cites.
Faster ridge regression via the subsampled randomized hadamard transform
Yichao Lu, Paramveer Dhillon, Dean P Foster, and Lyle Ungar · 2013
Earlier work this paper cites.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin · 2013
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Adaptively sparse transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 2019
Earlier work this paper cites.
Large memory layers with product keys
Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Earlier work this paper cites.
Mongoose: A learnable lsh framework for efficient neural network training
Beidi Chen, Zichang Liu, Binghui Peng, Zhaozhuo Xu, Jonathan Lingjie Li, Tri Dao, Zhao Song, Anshumali Shrivastava, and Christopher Re · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Earlier work this paper cites.
A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization
HanQin Cai, Yuchen Lou, Daniel Mckenzie, and Wotao Yin · 2021
Earlier work this paper cites.
Attention mechanism for neural machine translation: A survey
Weihua He, Yongyun Wu, and Xiaohua Li · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Earlier work this paper cites.
Sparse attention with learning to hash
Zhiqing Sun, Yiming Yang, and Shinjae Yoo · 2021
Earlier work this paper cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Earlier work this paper cites.
Optimizing language models for dialogue
ChatGPT · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Earlier work this paper cites.
Yichuan Deng, Wenyu Jin, Zhao Song, Xiaorui Sun, and Omri Weinstein · 2022
Earlier work this paper cites.
Attentive walk-aggregating graph neural networks
Mehmet F Demirel, Shengchao Liu, Siddhant Garg, Zhenmei Shi, and Yingyu Liang · 2022
Earlier work this paper cites.
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Cited alongside, same era.
Adore: Differentially oblivious relational database operators
Lianke Qin, Rajesh Jayaram, Elaine Shi, Zhao Song, Danyang Zhuo, and Shumo Chu · 2022
Cited alongside, same era.
Adaptive and dynamic multi-resolution hashing for pairwise summations
Lianke Qin, Aravind Reddy, Zhao Song, Zhaozhuo Xu, and Danyang Zhuo · 2022
Cited alongside, same era.
Deep online fused video stabilization
Zhenmei Shi, Fuhao Shi, Wei-Sheng Lai, Chia-Kai Liang, and Yingyu Liang · 2022
Cited alongside, same era.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2022
Cited alongside, same era.
Task-specific skill localization in fine-tuned language models
Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, and Sanjeev Arora · 2023
Later among the works it cites.
Fast submodular function maximization
Lianke Qin, Zhao Song, and Yitan Wang · 2023
Later among the works it cites.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2023
Later among the works it cites.
A general algorithm for solving rank-one matrix sensing
Lianke Qin, Zhao Song, and Ruizhe Zhang · 2023
Later among the works it cites.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Cited alongside, same era.
One pass streaming algorithm for super long token attention approximation in sublinear space
Raghav Addanki, Chenyang Li, Zhao Song, and Chiwun Yang · 2023
Cited alongside, same era.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al · 2023
Cited alongside, same era.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Cited alongside, same era.
Fine-tune language models to approximate unbiased in-context learning
Timothy Chu, Zhao Song, and Chiwun Yang · 2023
Cited alongside, same era.
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D.Manning, and Chelsea Finn · 2023
Later among the works it cites.
The trade-off between universality and label efficiency of representations from contrastive learning
Zhenmei Shi, Jiefeng Chen, Kunyang Li, Jayaram Raghuram, Xi Wu, Yingyu Liang, and Somesh Jha · 2023
Later among the works it cites.
When and how does known class help discover unknown ones? provable understanding through spectral analysis
Yiyou Sun, Zhenmei Shi, Yingyu Liang, and Yixuan Li · 2023
Later among the works it cites.
A unified scheme of resnet and softmax
Zhao Song, Weixin Wang, and Junze Yin · 2023
Later among the works it cites.
An automatic learning rate schedule algorithm for achieving faster convergence and steeper descent
Zhao Song and Chiwun Yang · 2023
Later among the works it cites.
Solving attention kernel regression problem via pre-conditioner
Zhao Song, Junze Yin, and Lichen Zhang · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Later among the works it cites.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Later among the works it cites.
The fine-grained complexity of gradient computation for training large language models
Josh Alman and Zhao Song · 2024
Closest in time.
How to capture higher-order correlations? generalizing matrix softmax attention to kronecker computation
Josh Alman and Zhao Song · 2024
Closest in time.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort · 2024
Closest in time.
How to protect copyright data in optimization of large language models?
Timothy Chu, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Jiuxiang Gu, Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2024
Closest in time.
Outlier-efficient hopfield layers for large transformer-based models
Jerry Yao-Chieh Hu, Pei-Hsuan Chang, Robin Luo, Hong-Yu Chen, Weijian Li, Wei-Po Wang, and Han Liu · 2024
Closest in time.
Nonparametric modern hopfield models
Jerry Yao-Chieh Hu, Bo-Yu Chen, Dennis Wu, Feng Ruan, and Han Liu · 2024
Closest in time.
On computational limits of modern hopfield models: A fine-grained complexity analysis
Jerry Yao-Chieh Hu, Thomas Lin, Zhao Song, and Han Liu · 2024
Closest in time.
Proxyformer: Nyström-based linear transformer with trainable proxy tokens
Sangho Lee, Hayun Lee, and Dongkun Shin · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Domain generalization via nuclear norm regularization
Zhenmei Shi, Yifei Ming, Ying Fan, Frederic Sala, and Yingyu Liang · 2024
Closest in time.
A graph-theoretic framework for understanding open-world semi-supervised learning
Yiyou Sun, Zhenmei Shi, and Yixuan Li · 2024
Closest in time.
Provable guarantees for neural networks via gradient feature learning
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Towards few-shot adaptation of foundation models via multitask finetuning
Zhuoyan Xu, Zhenmei Shi, Junyi Wei, Fangzhou Mu, Yin Li, and Yingyu Liang · 2024
Closest in time.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2024
Closest in time.