Fetching the paper…
Reading the bibliography…
Emerging Large Language Model (LLM) applications require long input context in order to perform complex tasks like document analysis and code generation.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
The fast multipole method for the wave equation: A pedestrian prescription
Ronald Coifman, Vladimir Rokhlin, and Stephen Wandzura. 1993 · 1993
Earlier work this paper cites.
N-body’problems in statistical learning
Alexander Gray and Andrew Moore. 2000 · 2000
Earlier work this paper cites.
Improved fast gauss transform and efficient kernel density estimation
Yang, Duraiswami, and Gumerov. 2003 · 2003
Earlier work this paper cites.
Dual-tree fast gauss transforms
Dongryeol Lee, Andrew Moore, and Alexander Gray. 2005 · 2005
Earlier work this paper cites.
Automatic online tuning for fast gaussian summation
Vlad Morariu, Balaji Srinivasan, Vikas C Raykar, Ramani Duraiswami, and Larry S Davis. 2008 · 2008
Earlier work this paper cites.
Askit: Approximate skeletonization kernel-independent treecode in high dimensions
William B March, Bo Xiao, and George Biros. 2015 · 2015
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling. arxiv preprint, 2019
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. 1911 · 2019
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and 1 others. 2020 · 2020
Earlier work this paper cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Earlier work this paper cites.
arxiv api user’s manual
2023
Earlier work this paper cites.
togethercomputer/llama-2-7b-32k
2023 · 2023
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023 · 2023
Earlier work this paper cites.
Claude 2: {https://www.anthropic.com/news/claude-2}
Anthropic. 2023 · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, and 1 others. 2023 · 2023
Cited alongside, same era.
Longlora: Efficient fine-tuning of long-context large language models
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2023 · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. 2023 · 2023
Cited alongside, same era.
Flash-decoding for long-context inference: {https://crfm.stanford.edu/2023/10/12/flashdecoding.html}
Tri Dao, Daniel Haziza, Francisco Massa, and Grigory Sisov. 2023 · 2023
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024 · 2024
Closest in time.
Ruler: What’s the real context size of your long-context language models?
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. 2024 · 2024
Closest in time.
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H Abdi, Dongsheng Li, Chin-Yew Lin, and 1 others. 2024 · 2024
Closest in time.
Gear: An efficient kv cache compression recipefor near-lossless generative inference of llm
Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, and Tuo Zhao. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. 2023 · 2023
Cited alongside, same era.
Fast multipole attention: A divide-and-conquer attention mechanism for long sequences
Yanming Kang, Giang Tran, and Hans De Sterck. 2023 · 2023
Cited alongside, same era.
How long can open-source llms truly promise on context length?
Dacheng Li, Rulin Shao Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. 2023 · 2023
Cited alongside, same era.
Fast attention over long sequences with dynamic sparse flash attention
Matteo Pagliardini, Daniele Paliotta, Martin Jaggi, and François Fleuret. 2023 · 2023
Cited alongside, same era.
Reducing transformer key-value cache size with cross-layer attention
William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, and Jonathan Ragan Kelly. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024 · 2024
Cited alongside, same era.
Lazyllm: Dynamic token pruning for efficient long context llm inference
Qichen Fu, Minsik Cho, Thomas Merth, Sachin Mehta, Mohammad Rastegari, and Mahyar Najibi. 2024 · 2024
Cited alongside, same era.
Gemini 1.5 https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024
Google. 2023 · 2024
Cited alongside, same era.
Minsoo Kim, Kyuhong Shim, Jungwook Choi, and Simyung Chang. 2024 · 2024
Closest in time.
Snapkv: Llm knows what you are looking for before generation
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. 2024 · 2024
Closest in time.
Dynamic memory compression: Retrofitting llms for accelerated inference
Piotr Nawrot, Adrian Łańcucki, Marcin Chochowski, David Tarjan, and Edoardo M Ponti. 2024 · 2024
Closest in time.
https://github.com/triton-lang/triton/blob/main/python/tutorials/06-fused-attention.py
OpenAI. 2024 · 2024
Closest in time.
Transformers are multi-state rnns
Matanel Oren, Michael Hassid, Yossi Adi, and Roy Schwartz. 2024 · 2024
Closest in time.
Quest: Query-aware sparsity for efficient long-context llm inference
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. 2024 · 2024
Closest in time.
Falcon 3 family of open foundation models
TII. 2024 · 2024
Closest in time.
Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference
Dongjie Yang, XiaoDong Han, Yan Gao, Yao Hu, Shilin Zhang, and Hai Zhao. 2024 · 2024
Closest in time.
Sirllm: Streaming infinite retentive llm
Yao Yao, Zuchao Li, and Hai Zhao. 2024 · 2024
Closest in time.