Fetching the paper…
Reading the bibliography…
Transformer-based large language models (LLMs) demonstrate impressive performance in long context generation.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Williams, S., Waterman, A., and Patterson, D · 2009
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Pipedream: Generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebron, F., and Sanghai, S · 2023
Earlier work this paper cites.
Half-quadratic quantization of large machine learning models, 2023
Badri, H. and Shaji, A · 2023
Earlier work this paper cites.
Extending context window of large language models via positional interpolation
Chen, S., Wong, S., Chen, L., and Tian, Y · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Earlier work this paper cites.
Longnet: Scaling transformers to 1,000,000,000 tokens
Ding, J., Ma, S., Dong, L., Zhang, X., Huang, S., Wang, W., Zheng, N., and Wei, F · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
Needle in a haystack-pressure testing llms
Kamradt, G · 2023
Cited alongside, same era.
Giraffe: Adventures in expanding context lengths in llms
Pal, A., Karkhanis, D., Roberts, M., Dooley, S., Sundararajan, A., and Naidu, S · 2023
Cited alongside, same era.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., et al · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Infinigen: Efficient generative inference of large language models with dynamic kv cache management
Lee, W., Lee, J., Seo, J., and Sim, J · 2024
Later among the works it cites.
Scbench: A kv cache-centric analysis of long-context methods
Li, Y., Jiang, H., Wu, Q., Luo, X., Ahn, S., Zhang, C., Abdi, A. H., Li, D., Gao, J., Yang, Y., et al · 2024
Later among the works it cites.
Loqt: Low-rank adapters for quantized pretraining
Loeschcke, S. B., Toftrup, M., Kastoryano, M., Belongie, S., and Snæbjarnarson, V · 2024
Later among the works it cites.
Mini-sequence transformers: Optimizing intermediate memory for long sequences training
Luo, C., Zhao, J., Chen, Z., Chen, B., and Anandkumar, A · 2024
Later among the works it cites.
Llama 3 gradient: A series of long context models, 2024
Pekelis, L., Feil, M., Moret, F., Huang, M., and Peng, T · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pose: Efficient context window extension of llms via positional skip-wise training
Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S · 2023
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B., Tumanov, A., and Ramjee, R · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, A · 2024
Cited alongside, same era.
Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks
Bai, Y., Tu, S., Zhang, J., Peng, H., Wang, X., Lv, X., Cao, S., Xu, J., Hou, L., Dong, Y., Tang, J., and Li, J · 2024
Cited alongside, same era.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
De, S., Smith, S. L., Fernando, A., Botev, A., Cristian-Muraru, G., Gu, A., Haroun, R., Berrada, L., Chen, Y., Srinivasan, S., et al · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Fu, Y., Cai, Z., Asi, A., Xiong, W., Dong, Y., and Xiao, W · 2024
Cited alongside, same era.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
Shadowkv: Kv cache in shadows for high-throughput long-context llm inference
Sun, H., Chang, L.-W., Bao, W., Zheng, S., Zheng, N., Liu, X., Dong, H., Chi, Y., and Chen, B · 2024
Later among the works it cites.
Razorattention: Efficient kv cache compression through retrieval heads
Tang, H., Lin, Y., Lin, J., Han, Q., Hong, S., Yao, Y., and Wang, G · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Later among the works it cites.
Retrieval head mechanistically explains long-context factuality
Wu, W., Wang, Y., Xiao, G., Peng, H., and Fu, Y · 2024
Later among the works it cites.
Layerkv: Optimizing large language model serving with layer-wise kv cache management
Xiong, Y., Wu, H., Shao, C., Wang, Z., Zhang, R., Guo, Y., Zhao, J., Zhang, K., and Pan, Z · 2024
Later among the works it cites.
vtensor: Flexible virtual tensor management for efficient llm serving
Xu, J., Zhang, R., Guo, C., Hu, W., Liu, Z., Wu, F., Feng, Y., Sun, S., Shao, C., Guo, Y., et al · 2024
Later among the works it cites.
Galore: Memory-efficient llm training by gradient low-rank projection
Zhao, J., Zhang, Z., Chen, B., Wang, Z., Anandkumar, A., and Tian, Y · 2024
Later among the works it cites.
A survey on efficient inference for large language models
Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., et al · 2024
Later among the works it cites.