Fetching the paper…
Reading the bibliography…
The self-attention mechanism is the key to the success of transformers in recent Large Language Models (LLMs).
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Algebraic complexity theory
Peter Bürgisser, Michael Clausen, and Mohammad A Shokrollahi · 2013
Earlier work this paper cites.
Fast matrix multiplication
Markus Bläser · 2013
Earlier work this paper cites.
Fast training of convolutional networks through ffts
Michael Mathieu, Mikael Henaff, and Yann LeCun · 2013
Earlier work this paper cites.
Super-resolution, extremal functions and the condition number of vandermonde matrices
Ankur Moitra · 2015
Earlier work this paper cites.
A robust sparse Fourier transform in the continuous setting
Eric Price and Zhao Song · 2015
Earlier work this paper cites.
Fourier-sparse interpolation without a frequency gap
Xue Chen, Daniel M Kane, Eric Price, and Zhao Song · 2016
Earlier work this paper cites.
Fast algorithms for convolutional neural networks
Andrew Lavin and Scott Gray · 2016
Earlier work this paper cites.
Recovery guarantee of weighted low-rank approximation via alternating minimization
Yuanzhi Li, Yingyu Liang, and Andrej Risteski · 2016
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P Woodruff · 2016
Earlier work this paper cites.
Fcnn: Fourier convolutional neural networks
Harry Pratt, Bryan Williams, Frans Coenen, and Yalin Zheng · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Earlier work this paper cites.
Sketching for kronecker product regression and p-splines
Huaian Diao, Zhao Song, Wen Sun, and David Woodruff · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Earlier work this paper cites.
Matrix Theory: Optimization, Concentration and Algorithms
Zhao Song · 2019
Earlier work this paper cites.
Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Fast fourier convolution
Lu Chi, Borui Jiang, and Yadong Mu · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Earlier work this paper cites.
Learning mixtures of linear regressions in subexponential time via fourier moments
Sitan Chen, Jerry Li, and Zhao Song · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2021
Earlier work this paper cites.
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Earlier work this paper cites.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Earlier work this paper cites.
An O ( k log n ) {O}(k\log n) time fourier set query algorithm
Yeqi Gao, Zhao Song, and Baocheng Sun · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Earlier work this paper cites.
Dynamic tensor product regression
Aravind Reddy, Zhao Song, and Lichen Zhang · 2022
Earlier work this paper cites.
Sparse fourier transform over lattices: A unified approach to signal reconstruction
Zhao Song, Baocheng Sun, Omri Weinstein, and Ruizhe Zhang · 2022
Earlier work this paper cites.
Linear complexity randomized self-attention mechanism
Lin Zheng, Chong Wang, and Lingpeng Kong · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Cited alongside, same era.
Longlora: Efficient fine-tuning of long-context large language models
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia · 2023
Cited alongside, same era.
Query complexity of active learning for function family with nearly orthogonal basis
Xiang Chen, Zhao Song, Baocheng Sun, Junze Yin, and Danyang Zhuo · 2023
Cited alongside, same era.
Superiority of softmax: Unveiling the performance edge over linear attention
Yichuan Deng, Zhao Song, and Tianyi Zhou · 2023
Cited alongside, same era.
Simple hardware-efficient long convolutions for sequence modeling
Daniel Y Fu, Elliot L Epstein, Eric Nguyen, Armin W Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré · 2023
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Closest in time.
The fine-grained complexity of gradient computation for training large language models
Josh Alman and Zhao Song · 2024
Closest in time.
How to capture higher-order correlations? generalizing matrix softmax attention to kronecker computation
Josh Alman and Zhao Song · 2024
Closest in time.
Unlimiformer: Long-range transformers with unlimited length input
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew Gormley · 2024
Closest in time.
Hsr-enhanced sparse attention acceleration
Bo Chen, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Cited alongside, same era.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Cited alongside, same era.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Cited alongside, same era.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Cited alongside, same era.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2023
Cited alongside, same era.
Super-resolution and robust sparse continuous fourier transform in any constant dimension: Nearly linear time and sample complexity
Yaonan Jin, Daogao Liu, and Zhao Song · 2023
Cited alongside, same era.
Ruisi Cai, Yuandong Tian, Zhangyang Wang, and Beidi Chen · 2024
Closest in time.
Longrope: Extending llm context window beyond 2 million tokens
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang · 2024
Closest in time.
Low rank matrix completion via robust alternating minimization in nearly linear time
Yuzhou Gu, Zhao Song, Junze Yin, and Lichen Zhang · 2024
Closest in time.
Outlier-efficient hopfield layers for large transformer-based models
Jerry Yao-Chieh Hu, Pei-Hsuan Chang, Haozheng Luo, Hong-Yu Chen, Weijian Li, Wei-Po Wang, and Han Liu · 2024
Closest in time.
Nonparametric modern hopfield models
Jerry Yao-Chieh Hu, Bo-Yu Chen, Dennis Wu, Feng Ruan, and Han Liu · 2024
Closest in time.
Hyperattention: Long-context attention in near-linear time
Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni, David Woodruff, and Amir Zandieh · 2024
Closest in time.
On computational limits of modern hopfield models: A fine-grained complexity analysis
Jerry Yao-Chieh Hu, Thomas Lin, Zhao Song, and Han Liu · 2024
Closest in time.
Provably optimal memory capacity for modern hopfield models: Tight analysis for transformer-compatible dense associative memories
Jerry Yao-Chieh Hu, Dennis Wu, and Han Liu · 2024
Closest in time.
Large language models for software engineering: A systematic literature review, 2024
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang · 2024
Closest in time.
Llm maybe longlm: Self-extend llm context window without tuning
Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Zirui Liu, Chia-Yuan Chang, Huiyuan Chen, and Xia Hu · 2024
Closest in time.
Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, et al · 2024
Closest in time.
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2024
Closest in time.
Beyond linear approximations: A novel pruning approach for attention matrix
Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Toward infinite-long prefix in transformer
Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
How to inverting the leverage score distribution?
Zhihang Li, Zhao Song, Weixin Wang, Junze Yin, and Zheng Yu · 2024
Closest in time.
Megalodon: Efficient llm pretraining and inference with unlimited context length
Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan May, Luke Zettlemoyer, Omer Levy, and Chunting Zhou · 2024
Closest in time.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D Lee · 2024
Closest in time.
Yarn: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Gautam Reddy · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Do large language models have compositional ability? an investigation into limitations and scalability
Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang · 2024
Closest in time.
Towards few-shot adaptation of foundation models via multitask finetuning
Zhuoyan Xu, Zhenmei Shi, Junyi Wei, Fangzhou Mu, Yin Li, and Yingyu Liang · 2024
Closest in time.
The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry
Michael Zhang, Kush Bhatia, Hermann Kumbong, and Christopher Re · 2024
Closest in time.