Fetching the paper…
Reading the bibliography…
There is growing demand for performing inference with hundreds of thousands of input tokens on trained transformer models.
Generating long sequences with sparse transformers, 2019
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text, 2016
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Malkov, Y. A. and Yashunin, D. A · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke, S., Ronan, L. B., Chandra, B., and Yejin, C · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Memory-efficient Transformers via Top-k Attention
Gupta, A., Dar, G., Goodman, S., Ciprut, D., and Berant, J · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Needle in a haystack - pressure testing llms
Kamradt, G · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Ring attention with blockwise transformers for near-infinite context, 2023
Liu, H., Zaharia, M., and Abbeel, P · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Later among the works it cites.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
Later among the works it cites.
Scaling rotational embeddings for long-context language models
Gradient Team · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
Cited alongside, same era.
Flexgen: high-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., Wang, Z., and Chen, B · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
AI, M · 2024
Cited alongside, same era.
Magicpig: Lsh sampling for efficient llm generation, 2024
Chen, Z., Sadhukhan, R., Ye, Z., Zhou, Y., Zhang, J., Nolte, N., Tian, Y., Douze, M., Bottou, L., Jia, Z., and Chen, B · 2024
Cited alongside, same era.
Later among the works it cites.
The llama 3 herd of models, 2024
Grattafiori, A. et al · 2024
Later among the works it cites.
The unreasonable ineffectiveness of the deeper layers, 2024
Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., and Roberts, D. A · 2024
Later among the works it cites.
RULER: What’s the Real Context Size of Your Long-Context Language Models?, April 2024
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., and Ginsburg, B · 2024
Later among the works it cites.
Loki: Low-rank keys for efficient sparse attention, 2024
Singhania, P., Singh, S., He, S., Feizi, S., and Bhatele, A · 2024
Later among the works it cites.
Quest: Query-aware sparsity for efficient long-context llm inference, 2024
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Later among the works it cites.
Pqcache: Product quantization-based kvcache for long context llm inference, 2024
Zhang, H., Ji, X., Chen, Y., Fu, F., Miao, X., Nie, X., Chen, W., and Cui, B · 2024
Later among the works it cites.