Fetching the paper…
Reading the bibliography…
This paper aims to overcome the "lost-in-the-middle" challenge of large language models (LLMs).
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2020
Earlier work this paper cites.
Booksum: A collection of datasets for long-form narrative summarization
Kryściński, W., Rajani, N., Agarwal, D., Xiong, C., and Radev, D · 2021
Earlier work this paper cites.
On the expressive power of self-attention matrices
Likhosherstov, V., Choromanski, K., and Weller, A · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Earlier work this paper cites.
Summˆ n: A multi-stage summarization framework for long input dialogues and documents
Zhang, Y., Ni, A., Mao, Z., Wu, C. H., Zhu, C., Deb, B., Awadallah, A. H., Radev, D., and Zhang, R · 2021
Earlier work this paper cites.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L., Goel, S., Kakade, S., and Zhang, C · 2022
Earlier work this paper cites.
Segnext: Rethinking convolutional attention design for semantic segmentation
Guo, M.-H., Lu, C.-Z., Hou, Q., Liu, Z., Cheng, M.-M., and Hu, S.-M · 2022
Earlier work this paper cites.
Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models
Varadi, M., Anyango, S., Deshpande, M., Nair, S., Natassia, C., Yordanova, G., Yuan, D., Stroe, O., Wood, G., Laydon, A., et al · 2022
Earlier work this paper cites.
Dialoglm: Pre-trained model for long dialogue understanding and summarization
Zhong, M., Liu, Y., Xu, Y., Zhu, C., and Zeng, M · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation
Du, X., Liu, M., Wang, K., Wang, H., Liu, J., Chen, Y., Feng, J., Sha, C., Peng, X., and Lou, Y · 2023
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
Team, M. N · 2023
Later among the works it cites.
Joma: Demystifying multilayer transformers via joint dynamics of mlp and attention
Tian, Y., Wang, Y., Zhang, Z., Chen, B., and Du, S · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Augmenting language models with long-term memory
Wang, W., Dong, L., Cheng, H., Liu, X., Yan, X., Gao, J., and Wei, F · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2023
Cited alongside, same era.
Lm-infinite: Simple on-the-fly length generalization for large language models
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Cited alongside, same era.
Efficient long-text understanding with short-text models
Ivgi, M., Shaham, U., and Berant, J · 2023
Cited alongside, same era.
Jacobs, S. A., Tanaka, M., Zhang, C., Zhang, M., Song, S. L., Rajbhandari, S., and He, Y · 2023
Cited alongside, same era.
Junqing, H., Kunhao, P., Xiaoqun, D., Zhuoyang, S., Yibo, L., Yuxin, L., Hao, W., Qianguo, S., Songxin, Z., Zejian, X., et al · 2023
Cited alongside, same era.
Loogle: Can long-context language models understand long contexts?
Li, J., Wang, M., Zheng, Z., and Zhang, M · 2023
Cited alongside, same era.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Cited alongside, same era.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2023
Cited alongside, same era.
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.
Retrieval meets long context large language models
Xu, P., Ping, W., Wu, X., McAfee, L., Zhu, C., Liu, Z., Subramanian, S., Bakhturina, E., Shoeybi, M., and Catanzaro, B · 2023
Later among the works it cites.
Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity, 2023
Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Pechenizkiy, M., Liang, Y., Wang, Z., and Liu, S · 2023
Later among the works it cites.
H _ 2 \_2 o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Later among the works it cites.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Wang, Z., Shen, L., Wang, A., Li, Y., Su, T., Yang, Z., and Tang, J · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Llm maybe longlm: Self-extend llm context window without tuning
Jin, H., Han, X., Yang, J., Jiang, Z., Liu, Z., Chang, C.-Y., Chen, H., and Hu, X · 2024
Closest in time.
Transformers are multi-state rnns
Oren, M., Hassid, M., Adi, Y., and Schwartz, R · 2024
Closest in time.
Can longer sequences help take the next leap in ai?, June 2022
Ré, C., Dao, T., Fu, D., and Goel, K · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
Soaring from 4k to 400k: Extending llm’s context with activation beacon
Zhang, P., Liu, Z., Xiao, S., Shao, N., Ye, Q., and Dou, Z · 2024
Closest in time.