Fetching the paper…
Reading the bibliography…
Leveraging long contexts is crucial for advanced AI systems, but attention computation poses a scalability challenge.
Similarity search in high dimensions via hashing
Gionis, A., Indyk, P., Motwani, R., et al · 1999
Earlier work this paper cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Shrivastava, A. and Li, P · 2014
Earlier work this paper cites.
On symmetric and asymmetric lshs for inner product search
Neyshabur, B. and Srebro, N · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Attention approximates sparse distributed memory
Bricken, T. and Pehlevan, C · 2021
Earlier work this paper cites.
Sparse attention with learning to hash
Sun, Z., Yang, Y., and Yoo, S · 2021
Earlier work this paper cites.
Zhang, A., Lipton, Z. C., Li, M., and Smola, A. J · 2021
Earlier work this paper cites.
Feng, J., Sun, Q., Xu, C., Zhao, P., Yang, Y., Tao, C., Zhao, D., and Lin, Q · 2022
Earlier work this paper cites.
LongT5: Efficient text-to-text transformer for long sequences
Guo, M., Ainslie, J., Uthus, D., Ontanon, S., Ni, J., Sung, Y.-H., and Yang, Y · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2023
Cited alongside, same era.
Hyperattention: Long-context attention in near-linear time
Han, I., Jayaram, R., Karbasi, A., Mirrokni, V., Woodruff, D. P., and Zandieh, A · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Using generative ai and gpt chatbots to improve instruction
Davies, R. and Murff, M · 2024
Closest in time.
Attention is naturally sparse with gaussian distributed input
Deng, Y., Song, Z., and Yang, C · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Gptfast, 2024
GPTFast · 2024
Closest in time.
Squeezed attention: Accelerating long context length llm inference
Hooper, C., Kim, S., Mohammadzadeh, H., Maheswaran, M., Paik, J., Mahoney, M. W., Keutzer, K., and Gholami, A · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Enhanced transformer with rotary position embedding., 2021
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Roformer, Y. L · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Cited alongside, same era.
Training-free long-context scaling of large language models
An, C., Huang, F., Zhang, J., Gong, S., Qiu, X., Zhou, C., and Kong, L · 2024
Cited alongside, same era.
LongBench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J · 2024
Cited alongside, same era.
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., and Ginsburg, B · 2024
Closest in time.
Retrieval augmented generation or long-context llms? a comprehensive study and hybrid approach
Li, Z., Li, C., Zhang, M., Mei, Q., and Bendersky, M · 2024
Closest in time.
Megalodon: Efficient llm pretraining and inference with unlimited context length
Ma, X., Yang, X., Xiong, W., Chen, B., Yu, L., Zhang, H., May, J., Zettlemoyer, L., Levy, O., and Zhou, C · 2024
Closest in time.
Loki: Low-rank keys for efficient sparse attention
Singhania, P., Singh, S., He, S., Feizi, S., and Bhatele, A · 2024
Closest in time.
Quest: Query-aware sparsity for efficient long-context llm inference
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Closest in time.
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., Han, S., and Sun, M · 2024
Closest in time.
Post-training sparse attention with double sparsity
Yang, S., Sheng, Y., Gonzalez, J. E., Stoica, I., and Zheng, L · 2024
Closest in time.
Flashinfer documentation, 2024
Ye, Z., Chen, L., Lai, R., Lin, W., Zhang, Y., Wang, S., Chen, T., Kasikci, B., Grover, V., Krishnamurthy, A., and Ceze, L · 2025
Closest in time.