Fetching the paper…
Reading the bibliography…
As the demand for long-context large language models (LLMs) increases, models with context windows of up to 128K or 1M tokens are becoming increasingly prevalent.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers, 2021
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M · 2021
Earlier work this paper cites.
Efficient attentions for long document summarization
Huang, L., Cao, S., Parulian, N., Ji, H., and Wang, L · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding, 2023
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
How long can open-source llms truly promise on context length?, June 2023
Li, D., Shao, R., Xie, A., Sheng, Y., Zheng, L., Gonzalez, J. E., Stoica, I., Ma, X., and Zhang, H · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models, 2023
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Cited alongside, same era.
Sparq attention: Bandwidth-efficient llm inference, 2023
Ribar, L., Chelombiev, I., Hudlass-Galley, L., Blake, C., Luschi, C., and Orr, D · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2023
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms, 2024
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Closest in time.
Nvidia ada lovelace professional gpu architecture
NVIDIA · 2024
Closest in time.
Nvbench: Nvidia’s benchmarking tool for gpus, 2024
NVIDIA · 2024
Closest in time.
New models and developer products announced at devday
OpenAI · 2024
Closest in time.
Introducing gpt-4o: our fastest and most affordable flagship model
OpenAI · 2024
Closest in time.
Transformers are multi-state RNNs, 2024
Oren, M., Hassid, M., Adi, Y., and Schwartz, R · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłoś, P · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
Introducing the next generation of Claude
Anthropic · 2024
Cited alongside, same era.
World model on million-length video and language with blockwise ringattention, 2024a
Liu, H., Yan, W., Zaharia, M., and Abbeel, P
Cited in the paper.
Scaling laws of rope-based extrapolation, 2024b
Liu, X., Yan, H., Zhang, S., An, C., Qiu, X., and Lin, D
Cited in the paper.
Parallel top-k algorithms on gpu: A comprehensive study and new methods
Zhang, J., Naruse, A., Li, X., and Wang, Y
Cited in the paper.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023b
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., Wang, Z., and Chen, B
Cited in the paper.
Cascade inference: Memory bandwidth efficient shared prefix batch decoding
Ye, Z., Lai, R., Lu, R., Lin, C.-Y., Zheng, S., Chen, L., Chen, T., and Ceze, L · 2024
Closest in time.
Atom: Low-bit quantization for efficient and accurate llm serving, 2024
Zhao, Y., Lin, C.-Y., Zhu, K., Ye, Z., Chen, L., Zheng, S., Ceze, L., Krishnamurthy, A., Chen, T., and Kasikci, B · 2024
Closest in time.