Fetching the paper…
Reading the bibliography…
Current Large Language Models (LLMs) are not only limited to some maximum context length, but also are not able to robustly consume long inputs.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
A theory of medical decision making and health: fuzzy trace theory
Reyna, V. F · 2008
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
A new intuitionism: Meaning, memory, and development in fuzzy-trace theory
Reyna, V. F · 2012
Earlier work this paper cites.
Reading wikipedia to answer open-domain questions, 2017
Chen, D., Fisch, A., Weston, J., and Bordes, A · 2017
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Kočiskỳ, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents, 2019
Dinan, E., Roller, S., Shuster, K., Fan, A., Auli, M., and Weston, J · 2019
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Interactive machine comprehension with information seeking agents
Yuan, X., Fu, J., Côté, M.-A., Tay, Y., Pal, C., and Trischler, A · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering, 2021
Izacard, G. and Grave, E · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Recursively summarizing books with human feedback
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P · 2021
Cited alongside, same era.
Qmsum: A new benchmark for query-based multi-domain meeting summarization
Zhong, M., Yin, D., Yu, T., Zaidi, A., Mutuma, M., Jha, R., Hassan, A., Celikyilmaz, A., Liu, Y., Qiu, X., et al · 2021
Cited alongside, same era.
LongT5: Efficient text-to-text transformer for long sequences
Guo, M., Ainslie, J., Uthus, D., Ontanon, S., Ni, J., Sung, Y.-H., and Yang, Y · 2022
Cited alongside, same era.
Quality: Question answering with long input texts, yes!
Pang, R. Y., Parrish, A., Joshi, N., Nangia, N., Phang, J., Chen, A., Padmakumar, V., Ma, J., Thompson, J., He, H., et al · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M · 2022
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2023
Later among the works it cites.
Lm-infinite: Simple on-the-fly length generalization for large language models
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Later among the works it cites.
Language models can solve computer tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Later among the works it cites.
Learning to reason and memorize with self-notes
Lanchantin, J., Toshniwal, S., Weston, J., Szlam, A., and Sukhbaatar, S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scrolls: Standardized comparison over long language sequences
Shaham, U., Segal, E., Ivgi, M., Efrat, A., Yoran, O., Haviv, A., Gupta, A., Xiong, W., Geva, M., Berant, J., et al · 2022
Cited alongside, same era.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2022
Cited alongside, same era.
Memorizing transformers
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Cited alongside, same era.
Re3: Generating longer stories with recursive reprompting and revision
Yang, K., Tian, Y., Peng, N., and Klein, D · 2022
Cited alongside, same era.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2022
Cited alongside, same era.
Colt5: Faster long-range transformers with conditional computation, 2023
Ainslie, J., Lei, T., de Jong, M., Ontañón, S., Brahma, S., Zemlyanskiy, Y., Uthus, D., Guo, M., Lee-Thorp, J., Tay, Y., Sung, Y.-H., and Sanghai, S · 2023
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Cited alongside, same era.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Later among the works it cites.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D · 2023
Later among the works it cites.
Pearl: Prompting large language models to plan and execute actions over long documents
Sun, S., Liu, Y., Wang, S., Zhu, C., and Iyyer, M · 2023
Later among the works it cites.
System 2 attention (is something you might need too)
Weston, J. and Sukhbaatar, S · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Memorybank: Enhancing large language models with long-term memory
Zhong, W., Guo, L., Gao, Q., and Wang, Y · 2023
Later among the works it cites.
Multimodal web navigation with instruction-finetuned foundation models
Furuta, H., Lee, K.-H., Nachum, O., Matsuo, Y., Faust, A., Gu, S. S., and Gur, I · 2024
Closest in time.
Llm maybe longlm: Self-extend llm context window without tuning
Jin, H., Han, X., Yang, J., Jiang, Z., Liu, Z., Chang, C.-Y., Chen, H., and Hu, X · 2024
Closest in time.