Fetching the paper…
Reading the bibliography…
Attention computation takes both the time complexity of $O(n^2)$ and the space complexity of $O(n^2)$ simultaneously, which makes deploying Large Language Models (LLMs) in streaming applications that involve long contexts requiring substantial computational resources.
Extensions of lipschitz mappings into a hilbert space
William B Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
The space complexity of approximating the frequency moments
Noga Alon, Yossi Matias, and Mario Szegedy · 1999
Earlier work this paper cites.
On graph problems in a semi-streaming model
Joan Feigenbaum, Sampath Kannan, Andrew McGregor, Siddharth Suri, and Jian Zhang · 2004
Earlier work this paper cites.
Finding graph matchings in data streams
Andrew McGregor · 2005
Earlier work this paper cites.
Hyperloglog: the analysis of a near-optimal cardinality estimation algorithm
Philippe Flajolet, Éric Fusy, Olivier Gandouet, and Frédéric Meunier · 2007
Earlier work this paper cites.
Linear programming in the semi-streaming model with application to the maximum matching problem
Kook Jin Ahn and Sudipto Guha · 2011
Earlier work this paper cites.
Bipartite matching in the semi-streaming model
Sebastian Eggert, Lasse Kliemann, Peter Munstermann, and Anand Srivastav · 2012
Earlier work this paper cites.
On the communication and streaming complexity of maximum bipartite matching
Ashish Goel, Michael Kapralov, and Sanjeev Khanna · 2012
Earlier work this paper cites.
Better bounds for matchings in the streaming model
Michael Kapralov · 2013
Earlier work this paper cites.
Economic efficiency requires interaction
Shahar Dobzinski, Noam Nisan, and Sigal Oren · 2014
Earlier work this paper cites.
Approximating matching size from random streams
Michael Kapralov, Sanjeev Khanna, and Madhu Sudan · 2014
Earlier work this paper cites.
A (2+ ϵ \epsilon )-approximation for maximum weight matching in the semi-streaming model
Ami Paz and Gregory Schwartzman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Access to data and number of iterations: Dual primal algorithms for maximum matching under resource constraints
Kook Jin Ahn and Sudipto Guha · 2018
Earlier work this paper cites.
Coresets meet edcs: algorithms for matching and vertex cover on massive graphs
Sepehr Assadi, MohammadHossein Bateni, Aaron Bernstein, Vahab Mirrokni, and Cliff Stein · 2019
Earlier work this paper cites.
Distributed and streaming linear programming in low dimensions
Sepehr Assadi, Nikolai Karpov, and Qin Zhang · 2019
Earlier work this paper cites.
Learning-based frequency estimation algorithms
Chen-Yu Hsu, Piotr Indyk, Dina Katabi, and Ali Vakilian · 2019
Earlier work this paper cites.
Heavy hitters via cluster-preserving clustering
Kasper Green Larsen, Jelani Nelson, Huy L Nguyen, and Mikkel Thorup · 2019
Earlier work this paper cites.
Stronger l2/l2 compressed sensing; without iterating
Vasileios Nakos and Zhao Song · 2019
Earlier work this paper cites.
Multi-pass graph streaming lower bounds for cycle counting, max-cut, matching size, and other problems
Sepehr Assadi, Gillat Kol, Raghuvansh R Saxena, and Huacheng Yu · 2020
Earlier work this paper cites.
Near-quadratic lower bounds for two-pass graph streaming algorithms
Sepehr Assadi and Ran Raz · 2020
Earlier work this paper cites.
Improved bound for matching in random-order streams
Aaron Bernstein · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Streaming complexity of spanning tree computation
Yi-Jun Chang, Martin Farach-Colton, Tsan-Sheng Hsu, and Meng-Tsung Tsai · 2020
Earlier work this paper cites.
Approximate maximum matching in random streams
Alireza Farhadi, Mohammad Taghi Hajiaghayi, Tung Mah, Anup Rao, and Ryan A Rossi · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Earlier work this paper cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Earlier work this paper cites.
An auction algorithm for bipartite matching in streaming and massively parallel computation models
Sepehr Assadi, S Cliff Liu, and Robert E Tarjan · 2021
Earlier work this paper cites.
Mongoose: A learnable lsh framework for efficient neural network training
Beidi Chen, Zichang Liu, Binghui Peng, Zhaozhuo Xu, Jonathan Lingjie Li, Tri Dao, Zhao Song, Anshumali Shrivastava, and Christopher Re · 2021
Cited alongside, same era.
Perfect l_p sampling in a data stream
Rajesh Jayaram and David Woodruff · 2021
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
Synthesizer: Rethinking self-attention for transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2021
Cited alongside, same era.
Sgcsumm: An extractive multi-document summarization method based on pre-trained language model, submodularity, and graph convolutional neural networks
Alireza Ghadimi and Hamid Beigy · 2023
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Closest in time.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semi-streaming bipartite matching in fewer passes and optimal space
Sepehr Assadi, Arun Jambulapati, Yujia Jin, Aaron Sidford, and Kevin Tian · 2022
Cited alongside, same era.
Optimizing language models for dialogue
ChatGPT · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Cited alongside, same era.
Adobe firefly
Adobe · 2023
Cited alongside, same era.
Insu Han, Rajesh Jarayam, Amin Karbasi, Vahab Mirrokni, David P Woodruff, and Amir Zandieh · 2023
Closest in time.
Kung-Hsiang Huang, Philippe Laban, Alexander R Fabbri, Prafulla Kumar Choubey, Shafiq Joty, Caiming Xiong, and Chien-Sheng Wu · 2023
Closest in time.
Longeval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo · 2023
Closest in time.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Praneeth Kacham, Vahab Mirrokni, and Peilin Zhong · 2023
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
How long can opensource llms truly promise on context length, 2023
Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph E Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Space-efficient interior point method, with applications to linear programming and maximum weight bipartite matching
S Cliff Liu, Zhao Song, Hengjie Zhang, Lichen Zhang, and Tianyi Zhou · 2023
Closest in time.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Closest in time.
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang · 2023
Closest in time.
Recent advances in deep learning based dialogue systems: A systematic survey
Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, and Erik Cambria · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Trainable transformer in transformer
Abhishek Panigrahi, Sadhika Malladi, Mengzhou Xia, and Sanjeev Arora · 2023
Closest in time.
Yarn: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2023
Closest in time.
Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension
Anna Rogers, Matt Gardner, and Isabelle Augenstein · 2023
Closest in time.
Analysis of community question-answering issues via machine learning and deep learning: State-of-the-art review
Pradeep Kumar Roy, Sunil Saumya, Jyoti Prakash Singh, Snehasish Banerjee, and Adnan Gutub · 2023
Closest in time.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Closest in time.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Transformers as support vector machines
Davoud Ataee Tarzanagh, Yingcong Li, Christos Thrampoulidis, and Samet Oymak · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Sheared llama: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.
Extractive summarization via chatgpt for faithful summary generation
Haopeng Zhang, Xiao Liu, and Jiawei Zhang · 2023
Closest in time.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, et al · 2023
Closest in time.