Fetching the paper…
Reading the bibliography…
Long prompt leads to huge hardware costs when using transformer-based Large Language Models (LLMs).
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q.V., Salakhutdinov, R., 2019 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., Sutskever, I., 2019 · 1904
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M.E., Cohan, A., 2020 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, pp. 74–81
Lin, C.Y., 2004 · 2004
Earlier work this paper cites.
Leveraging duc, in: proceedings of DUC
Copeck, T., Inkpen, D., Kazantseva, A., Kennedy, A., Kipp, D., Nastase, V., Szpakowicz, S., 2006 · 2006
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al., 2020 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al., 2020 · 2010
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Cohan, A., Dernoncourt, F., Kim, D.S., Bui, T., Kim, S., Chang, W., Goharian, N., 2018 · 2018
Earlier work this paper cites.
Narayan, S., Cohen, S.B., Lapata, M., 2018 · 2018
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention, in: International conference on machine learning, PMLR. pp. 5156–5165
Katharopoulos, A., Vyas, A., Pappas, N., Fleuret, F., 2020 · 2020
Earlier work this paper cites.
Fnet: Mixing tokens with fourier transforms
Lee-Thorp, J., Ainslie, J., Eckstein, I., Ontanon, S., 2021 · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X.L., Liang, P., 2021 · 2021
Earlier work this paper cites.
Token merging: Your vit but faster
Bolya, D., Fu, C.Y., Dai, X., Zhang, P., Feichtenhofer, C., Hoffman, J., 2022 · 2022
Earlier work this paper cites.
Recurrent memory transformer
Bulatov, A., Kuratov, Y., Burtsev, M., 2022 · 2022
Earlier work this paper cites.
Cicero: A dataset for contextualized commonsense inference in dialogues
Ghosal, D., Shen, S., Majumder, N., Mihalcea, R., Poria, S., 2022 · 2022
Cited alongside, same era.
Learned token pruning for transformers, in: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 784–794
Kim, S., Shen, S., Thorsley, D., Gholami, A., Kwon, W., Hassoun, J., Keutzer, K., 2022 · 2022
Cited alongside, same era.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 61–68
Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., Tang, J., 2022 · 2022
Cited alongside, same era.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A.S., Naik, A., Stap, D., et al., 2022 · 2022
Cited alongside, same era.
In-context autoencoder for context compression in a large language model
Ge, T., Hu, J., Wang, X., Chen, S.Q., Wei, F., 2023 · 2023
Later among the works it cites.
Llmlingua: Compressing prompts for accelerated inference of large language models
Jiang, H., Wu, Q., Lin, C.Y., Yang, Y., Qiu, L., 2023 · 2023
Later among the works it cites.
Li, Y., 2023 · 2023
Later among the works it cites.
Learning to compress prompts with gist tokens
Mu, J., Li, X.L., Goodman, N., 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., Hashimoto, T.B., 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wingate, D., Shoeybi, M., Sorensen, T., 2022 · 2022
Cited alongside, same era.
Wu, Y., Rabe, M.N., Hutchins, D., Szegedy, C., 2022 · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X.V., et al., 2022 · 2022
Cited alongside, same era.
Linear complexity randomized self-attention mechanism, in: International conference on machine learning, PMLR. pp. 27011–27041
Zheng, L., Wang, C., Kong, L., 2022 · 2022
Cited alongside, same era.
Training language models with memory augmentation
Zhong, Z., Lei, T., Chen, D., 2022 · 2022
Cited alongside, same era.
Unlimiformer: Long-range transformers with unlimited length input
Bertsch, A., Alon, U., Neubig, G., Gormley, M.R., 2023 · 2023
Cited alongside, same era.
Scaling transformer to 1m tokens and beyond with rmt
Bulatov, A., Kuratov, Y., Kapushev, Y., Burtsev, M.S., 2023 · 2023
Cited alongside, same era.
Adapting language models to compress contexts
Chevalier, A., Wettig, A., Ajith, A., Chen, D., 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Bluelm: An open multilingual 7b language model
Team, B., 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al., 2023 · 2023
Later among the works it cites.
Zero-shot information extraction via chatting with chatgpt
Wei, X., Cui, X., Cheng, N., Wang, X., Zhang, X., Huang, S., Xie, P., Xu, J., Chen, Y., Zhang, M., et al., 2023 · 2023
Later among the works it cites.
Exploring the limits of chatgpt for query or aspect-based text summarization
Yang, X., Li, Y., Zhang, X., Chen, H., Cheng, W., 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al., 2023 · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P., 2024 · 2024
Closest in time.
Dialogue acts enhanced extract–abstract framework for meeting summarization
Sun, S., Yuan, R., Li, W., Cao, Z., Li, S., 2024 · 2024
Closest in time.
Dialogue summarization enhanced response generation for multi-domain task-oriented dialogue systems
Wang, L., Zhao, M., Ji, H., Jiang, Z., Li, R., Hu, Z., Lu, X., 2024 · 2024
Closest in time.