Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) exhibit exceptional proficiency in handling extensive context windows in natural language.
Gaussian elimination is not optimal
Volker Strassen · 1969
Earlier work this paper cites.
I/o complexity: The red-blue pebble game
Jia-Wei Hong and Hsiang-Tsung Kung · 1981
Earlier work this paper cites.
Matrix multiplication via arithmetic progressions
Don Coppersmith and Shmuel Winograd · 1987
Earlier work this paper cites.
The input/output complexity of sorting and related problems
Alok Aggarwal and S Vitter, Jeffrey · 1988
Earlier work this paper cites.
External memory algorithms and data structures: Dealing with massive data
Jeffrey Scott Vitter · 2001
Earlier work this paper cites.
Using advanced MPI: Modern features of the message-passing interface
William Gropp, Torsten Hoefler, Rajeev Thakur, and Ewing Lusk · 2014
Earlier work this paper cites.
The input/output complexity of sparse matrix multiplication
Rasmus Pagh and Morten Stöckel · 2014
Earlier work this paper cites.
Minimax rates for memory-bounded sparse linear regression
Jacob Steinhardt and John Duchi · 2015
Earlier work this paper cites.
In-memory big data management and processing: A survey
Hao Zhang, Gang Chen, Beng Chin Ooi, Kian-Lee Tan, and Meihui Zhang · 2015
Earlier work this paper cites.
The i/o complexity of computing prime tables
Michael A Bender, Rezaul Chowdhury, Alexander Conway, Martin Farach-Colton, Pramod Ganapathi, Rob Johnson, Samuel McCauley, Bertrand Simon, and Shikha Singh · 2016
Earlier work this paper cites.
Memory, communication, and statistical queries
Jacob Steinhardt, Gregory Valiant, and Stefan Wager · 2016
Earlier work this paper cites.
On the space complexity of linear programming with preprocessing
Yael Tauman Kalai, Ran Raz, and Oded Regev · 2016
Earlier work this paper cites.
Erik D Demaine, Andrea Lincoln, Quanquan C Liu, Jayson Lynch, and Virginia Vassilevska Williams · 2017
Earlier work this paper cites.
A general memory-bounded learning algorithm
Michal Moshkovitz and Naftali Tishby · 2017
Earlier work this paper cites.
A time-space lower bound for a large class of learning problems
Ran Raz · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Red-blue pebble game: Complexity of computing the trade-off between cache size and memory transfers
Erik D Demaine and Quanquan C Liu · 2018
Earlier work this paper cites.
Extractor-based time-space lower bounds for learning
Sumegha Garg, Ran Raz, and Avishay Tal · 2018
Earlier work this paper cites.
Fast learning requires good memory: A time-space lower bound for parity learning
Ran Raz · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit · 2018
Earlier work this paper cites.
Estimating entropy of distributions in constant space
Jayadev Acharya, Sourbh Bhadane, Piotr Indyk, and Ziteng Sun · 2019
Earlier work this paper cites.
The i/o complexity of toom-cook integer multiplication
Gianfranco Bilardi and Lorenzo De Stefani · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
The i/o complexity of hybrid algorithms for square matrix multiplication
Lorenzo De Stefani · 2019
Earlier work this paper cites.
On the i/o complexity of hybrid algorithms for integer multiplication
Lorenzo De Stefani · 2019
Earlier work this paper cites.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric Villemonte de La Clergerie, Djamé Seddah, and Benoit Sagot · 2019
Earlier work this paper cites.
Revisiting the i/o-complexity of fast matrix multiplication with recomputations
Roy Nissim and Oded Schwartz · 2019
Earlier work this paper cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Earlier work this paper cites.
Memory-sample tradeoffs for linear regression with small error
Vatsal Sharan, Aaron Sidford, and Gregory Valiant · 2019
Earlier work this paper cites.
Open problem: The oracle complexity of convex optimization with limited memory
Blake Woodworth and Nathan Srebro · 2019
Earlier work this paper cites.
Bp-transformer: Modelling long-range context via binary partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Improving i/o complexity of triangle enumeration
Yi Cui, Di Xiao, Daren BH Cline, and Dmitri Loguinov · 2020
Earlier work this paper cites.
Towards a combinatorial characterization of bounded-memory learning
Alon Gonen, Shachar Lovett, and Michal Moshkovitz · 2020
Earlier work this paper cites.
Spectral lower bounds on the i/o complexity of computation graphs
Saachi Jain and Matei Zaharia · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
S Cliff Liu, Zhao Song, Hengjie Zhang, Lichen Zhang, and Tianyi Zhou · 2020
Earlier work this paper cites.
Sparse sinkhorn attention
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan · 2020
Earlier work this paper cites.
Attention-based sentiment analysis using convolutional and recurrent neural network
Mohd Usama, Belal Ahmad, Enmin Song, M Shamim Hossain, Mubarak Alrashoud, and Ghulam Muhammad · 2020
Earlier work this paper cites.
O (n) connections are expressive enough: Universal approximability of sparse transformers
Chulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Earlier work this paper cites.
Bounded memory active learning through enriched queries
Max Hopkins, Daniel Kane, Shachar Lovett, and Michal Moshkovitz · 2021
Earlier work this paper cites.
Attention mechanism for neural machine translation: a survey
Weihua He, Yongyun Wu, and Xiaohua Li · 2021
Earlier work this paper cites.
I/o efficient k-truss community search in massive graphs
Yuli Jiang, Xin Huang, and Hong Cheng · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Cited alongside, same era.
Multi-armed bandits with bounded arm-memory: Near-optimal guarantees for best-arm identification and regret minimization
Arnab Maiti, Vishakha Patil, and Arindam Khan · 2021
Cited alongside, same era.
Optimal-degree polynomial approximations for exponentials and gaussian kernel density estimation
Amol Aggarwal and Josh Alman · 2022
Cited alongside, same era.
Estimation of entropy in constant space with improved sample complexity
Maryam Aliakbarpour, Andrew McGregor, Jelani Nelson, and Erik Waingarten · 2022
Cited alongside, same era.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2024
Closest in time.
Internet of agents: Weaving a web of heterogeneous agents for collaborative intelligence
Weize Chen, Ziming You, Ran Li, Yitong Guan, Chen Qian, Chenyang Zhao, Cheng Yang, Ruobing Xie, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Tri Dao and Albert Gu · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et al · 2024
Closest in time.
Attention is naturally sparse with gaussian distributed input
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Strong memory lower bounds for learning natural models
Gavin Brown, Mark Bun, and Adam Smith · 2022
Cited alongside, same era.
Memory bounds for continual learning
Xi Chen, Christos Papadimitriou, and Binghui Peng · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Memory efficient continual learning with transformers
Beyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal, and Cedric Archambeau · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Yichuan Deng, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Subgraph enumeration in optimal i/o complexity
Shiyuan Deng and Yufei Tao · 2024
Closest in time.
Tao Feng, Chuanyang Jin, Jingyu Liu, Kunlun Zhu, Haoqin Tu, Zirui Cheng, Guanyu Lin, and Jiaxuan You · 2024
Closest in time.
Leo Feng, Frederick Tung, Hossein Hajimirsadeghi, Mohamed Osama Ahmed, Yoshua Bengio, and Greg Mori · 2024
Closest in time.
Outlier-efficient hopfield layers for large transformer-based models
Jerry Yao-Chieh Hu, Pei-Hsuan Chang, Haozheng Luo, Hong-Yu Chen, Weijian Li, Wei-Po Wang, and Han Liu · 2024
Closest in time.
Nonparametric modern hopfield models
Jerry Yao-Chieh Hu, Bo-Yu Chen, Dennis Wu, Feng Ruan, and Han Liu · 2024
Closest in time.
Hyperattention: Long-context attention in near-linear time
Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni, David Woodruff, and Amir Zandieh · 2024
Closest in time.
On computational limits of modern hopfield models: A fine-grained complexity analysis
Jerry Yao-Chieh Hu, Thomas Lin, Zhao Song, and Han Liu · 2024
Closest in time.
Computational limits of low-rank adaptation (lora) for transformer-based models
Jerry Yao-Chieh Hu, Maojiang Su, En-Jui Kuo, Zhao Song, and Han Liu · 2024
Closest in time.
Wikibench: Community-driven data curation for ai evaluation on wikipedia
Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu · 2024
Closest in time.
Na Liu, Liangyu Chen, Xiaoyu Tian, Wei Zou, Kaijiang Chen, and Ming Cui · 2024
Closest in time.
Yingyu Liang, Heshan Liu, Zhenmei Shi, Zhao Song, Zhuoyan Xu, and Junze Yin · 2024
Closest in time.
Beyond linear approximations: A novel pruning approach for attention matrix, 2024
Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Toward infinite-long prefix in transformer
Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
AI @ Meta Llama Team · 2024
Closest in time.
Evaluating very long-term conversational memory of llm agents
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang · 2024
Closest in time.
Linearizing large language models
Jean Mercat, Igor Vasiljevic, Sedrick Keh, Kushal Arora, Achal Dave, Adrien Gaidon, and Thomas Kollar · 2024
Closest in time.
Introducing openai o1-preview
OpenAI · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Whispered tuning: Data privacy preservation in fine-tuning llms through differential privacy
Tanmay Singh, Harshvardhan Aditya, Vijay K Madisetti, and Arshdeep Bahga · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Closest in time.
Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding
Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
Why larger language models do in-context learning differently?
Zhenmei Shi, Junyi Wei, Zhuoyan Xu, and Yingyu Liang · 2024
Closest in time.
I/o complexity of attention, or how optimal is flashattention?
Barna Saha and Christopher Ye · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Is a picture worth a thousand words? delving into spatial reasoning for vision language models
Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, and Neel Joshi · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Do large language models have compositional ability? an investigation into limitations and scalability
Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang · 2024
Closest in time.
The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry
Michael Zhang, Kush Bhatia, Hermann Kumbong, and Christopher Ré · 2024
Closest in time.
Vision-language models for vision tasks: A survey
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu · 2024
Closest in time.
Self-guide: Better task-specific instruction following via self-synthetic finetuning
Chenyang Zhao, Xueying Jia, Vijay Viswanathan, Tongshuang Wu, and Graham Neubig · 2024
Closest in time.
The expressive power of low-rank adaptation
Yuchen Zeng and Kangwook Lee · 2024
Closest in time.