Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities across various applications, but their performance on long-context tasks is often limited by the computational complexity of attention mechanisms.
On a modification of chebyshev’s inequality and of the error formula of laplace
Sergei Bernstein · 1924
Earlier work this paper cites.
Dynamic half-space reporting, geometric optimization, and minimum spanning trees
Pankaj K Agarwal, David Eppstein, and Jirí Matousek · 1992
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Geometric algorithms for density-based data clustering
Danny Z Chen, Michiel Smid, and Bin Xu · 2005
Earlier work this paper cites.
On some geometric problems of color-spanning sets
Wenqi Ju, Chenglin Fan, Jun Luo, Binhai Zhu, and Ovidiu Daescu · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Continuously differentiable exponential linear units
Jonathan T Barron · 2017
Earlier work this paper cites.
Defining equitable geographic districts in road networks via stable matching
David Eppstein, Michael T Goodrich, Doruk Korkmaz, and Nil Mamano · 2017
Earlier work this paper cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Wordcraft: A human-ai collaborative editor for story writing
Andy Coenen, Luke Davis, Daphne Ippolito, Emily Reif, and Ann Yuan · 2021
Earlier work this paper cites.
A faster algorithm for solving general lps
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2021
Earlier work this paper cites.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Earlier work this paper cites.
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Earlier work this paper cites.
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le · 2022
Earlier work this paper cites.
Robust training of neural networks using scale invariant architectures
Zhiyuan Li, Srinadh Bhojanapalli, Manzil Zaheer, Sashank Reddi, and Sanjiv Kumar · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Earlier work this paper cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Earlier work this paper cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Earlier work this paper cites.
Dynamic context pruning for efficient and interpretable autoregressive transformers
Sotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci, Aurelien Lucchi, and Thomas Hofmann · 2023
Earlier work this paper cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Earlier work this paper cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Earlier work this paper cites.
Dynamic algorithms for packing-covering lps via multiplicative weight updates
Sayan Bhattacharya, Peter Kiss, and Thatchaphol Saranurak · 2023
Earlier work this paper cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Earlier work this paper cites.
What can a single attention layer learn? a study through the random features lens
Hengyu Fu, Tianyu Guo, Yu Bai, and Song Mei · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Earlier work this paper cites.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Earlier work this paper cites.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2023
Earlier work this paper cites.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Praneeth Kacham, Vahab Mirrokni, and Peilin Zhong · 2023
Earlier work this paper cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Earlier work this paper cites.
Gpt-4 turbo, 2023
OpenAI · 2023
Earlier work this paper cites.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2023
Cited alongside, same era.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2023
Cited alongside, same era.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Cited alongside, same era.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Toward infinite-long prefix in transformer
Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Unraveling the smoothness properties of diffusion models: A gaussian mixture perspective
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kai Shen, Junliang Guo, Xu Tan, Siliang Tang, Rui Wang, and Jiang Bian · 2023
Cited alongside, same era.
Sketching meets differential privacy: fast algorithm for dynamic kronecker projection maintenance
Zhao Song, Xin Yang, Yuanyuan Yang, and Lichen Zhang · 2023
Cited alongside, same era.
Streaming semidefinite programs: O ( n ) {O}(\sqrt{n}) passes, small space and fast runtime
Zhao Song, Mingquan Ye, and Lichen Zhang · 2023
Cited alongside, same era.
Dolfin: Diffusion layout transformers without autoencoder
Yilin Wang, Zeyuan Chen, Liangjun Zhong, Zheng Ding, Zhizhou Sha, and Zhuowen Tu · 2023
Cited alongside, same era.
Replacing softmax with relu in vision transformers
Mitchell Wortsman, Jaehoon Lee, Justin Gilmer, and Simon Kornblith · 2023
Cited alongside, same era.
Tokencompose: Grounding diffusion with token-level supervision
Zirui Wang, Zhizhou Sha, Zheng Ding, Yilin Wang, and Zhuowen Tu · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Cited alongside, same era.
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Llama 3, 2024
Meta · 2024
Closest in time.
Mistral nemo, 2024
MistralAI · 2024
Closest in time.
Linearizing large language models
Jean Mercat, Igor Vasiljevic, Sedrick Keh, Kushal Arora, Achal Dave, Adrien Gaidon, and Thomas Kollar · 2024
Closest in time.
Massive activations in large language models
Mingjie Sun, Xinlei Chen, J Zico Kolter, and Zhuang Liu · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
Lazydit: Lazy learning for the acceleration of diffusion transformers
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Yanyu Li, Yifan Gong, Kai Zhang, Hao Tan, Jason Kuen, Henghui Ding, et al · 2024
Closest in time.
Numerical pruning for efficient autoregressive models
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Jing Liu, Ruiyi Zhang, Ryan A Rossi, Hao Tan, Tong Yu, Xiang Chen, et al · 2024
Closest in time.
Why larger language models do in-context learning differently?
Zhenmei Shi, Junyi Wei, Zhuoyan Xu, and Yingyu Liang · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.
Quest: Query-aware sparsity for efficient long-context llm inference
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Is a picture worth a thousand words? delving into spatial reasoning for vision language models
Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Yixuan Li, and Neel Joshi · 2024
Closest in time.
Transformers are deep optimizers: Provable in-context learning for deep model training
Weimin Wu, Maojiang Su, Jerry Yao-Chieh Hu, Zhao Song, and Han Liu · 2024
Closest in time.
Omnicontrolnet: Dual-stage integration for conditional image generation
Yilin Wang, Haiyang Xu, Xiang Zhang, Zeyuan Chen, Zhizhou Sha, Zirui Wang, and Zhuowen Tu · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Do large language models have compositional ability? an investigation into limitations and scalability
Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang · 2024
Closest in time.
The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry
Michael Zhang, Kush Bhatia, Hermann Kumbong, and Christopher Ré · 2024
Closest in time.
Subgen: Token generation in sublinear time and memory
Amir Zandieh, Insu Han, Vahab Mirrokni, and Amin Karbasi · 2024
Closest in time.
Benchmarking large language models for news summarization
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto · 2024
Closest in time.
Yang Cao, Bo Chen, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Mingda Wan · 2025
Closest in time.
High-order matching for one-step shortcut diffusion models
Bo Chen, Chengyue Gong, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Mingda Wan · 2025
Closest in time.
Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent
Bo Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Nrflow: Towards noise-robust generative modeling via second-order flow matching
Bo Chen, Xiaoyu Li, Yingyu Liang, Zhao Song, and Zhizhou Sha · 2025
Closest in time.
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Curse of attention: A kernel-based perspective for why transformers fail to generalize on time series forecasting and beyond
Yekun Ke, Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2025
Closest in time.
Neural algorithmic reasoning for hypergraphs with looped transformers
Xiaoyu Li, Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Zhen Zhuang · 2025
Closest in time.
Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2025
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2025
Closest in time.