Fetching the paper…
Reading the bibliography…
In recent years, the input context sizes of large language models (LLMs) have increased dramatically.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, J., Bordes, A., Chopra, S., and Mikolov, T · 2016
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Legendre memory units: Continuous-time representation in recurrent neural networks
Voelker, A., Kajić, I., and Eliasmith, C · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Transformers as soft reasoners over language
Clark, P., Tafjord, O., and Richardson, K · 2020
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Fan, A., Lavril, T., Grave, E., Joulin, A., and Sukhbaatar, S · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning, 2020
Lei, J., Wang, L., Shen, Y., Yu, D., Berg, T. L., and Bansal, M · 2020
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Earlier work this paper cites.
Datasheets for datasets
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Iii, H. D., and Crawford, K · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Re, C · 2021
Earlier work this paper cites.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Earlier work this paper cites.
ProofWriter: Generating implications, proofs, and abductive statements over natural language
Tafjord, O., Dalvi, B., and Clark, P · 2021
Earlier work this paper cites.
Long range arena : A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Earlier work this paper cites.
Diagnosing the first-order logical reasoning ability through LogicNLI
Tian, J., Li, Y., Chen, W., Xiao, L., He, H., and Jin, Y · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2022
Earlier work this paper cites.
Recurrent memory transformer
Bulatov, A., Kuratov, Y., and Burtsev, M · 2022
Earlier work this paper cites.
LangChain, October 2022
Chase, H · 2022
Cited alongside, same era.
Temporal latent bottleneck: Synthesis of fast and slow processing mechanisms in sequence learning
Didolkar, A., Gupta, K., Goyal, A., Gundavarapu, N. B., Lamb, A. M., Ke, N. R., and Bengio, Y · 2022
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Cited alongside, same era.
FOLIO: natural language reasoning with first-order logic
Han, S., Schoelkopf, H., Zhao, Y., Qi, Z., Riddell, M., Benson, L., Sun, L., Zubova, E., Qiao, Y., Burtell, M., Peng, D., Fan, J., Liu, Y., Wong, B., Sailor, M., Ni, A., Nan, L., Kasai, J., Yu, T., Zhang, R., Joty, S. R., Fabbri, A. R., Kryscinski, W., Lin, X. V., Xiong, C., and Radev, D · 2022
Cited alongside, same era.
Block-recurrent transformers
Hutchins, D., Schlag, I., Wu, Y., Dyer, E., and Neyshabur, B · 2022
Cited alongside, same era.
Marathon: A race through the realm of long context with large language models
Zhang, L., Li, Y., Liu, Z., Liu, J., Yang, M., et al · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M., Jacobs, S. A., Awan, A. A., Aneja, J., Awadallah, A., Awadalla, H., Bach, N., Bahree, A., Bakhtiari, A., Behl, H., et al · 2024
Closest in time.
Agarwal, R., Singh, A., Zhang, L. M., Bohnet, B., Chan, S., Anand, A., Abbas, Z., Nova, A., Co-Reyes, J. D., Chu, E., et al · 2024
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Introducing the next generation of claude, 2024
Anthropic · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scrolls: Standardized comparison over long language sequences
Shaham, U., Segal, E., Ivgi, M., Efrat, A., Yoran, O., Haviv, A., Gupta, A., Xiong, W., Geva, M., Berant, J., et al · 2022
Cited alongside, same era.
Explain my surprise: Learning efficient long-term memory by predicting uncertain outcomes
Sorokin, A., Buzun, N., Pugachev, L., and Burtsev, M · 2022
Cited alongside, same era.
Chapterbreak: A challenge dataset for long-range language models
Sun, S., Thai, K., and Iyyer, M · 2022
Cited alongside, same era.
Memformer: A memory-augmented transformer for sequence modeling
Wu, Q., Lan, Z., Qian, K., Gu, J., Geramifard, A., and Yu, Z · 2022
Cited alongside, same era.
L-eval: Instituting standardized evaluation for long context language models
An, C., Gong, S., Zhong, M., Li, M., Zhang, J., Kong, L., and Qiu, X · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Cited alongside, same era.
Longlora: Efficient fine-tuning of long-context large language models
Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J · 2023
Cited alongside, same era.
Bai, Y., Lv, X., Zhang, J., He, Y., Qi, J., Hou, L., Tang, J., Dong, Y., and Li, J · 2024
Closest in time.
Beyond attention: Breaking the limits of transformer context length with recurrent memory
Bulatov, A., Kuratov, Y., Kapushev, Y., and Burtsev, M · 2024
Closest in time.
Command r: Retrieval-augmented generation at production scale, March 2024
Cohere · 2024
Closest in time.
The faiss library
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P.-E., Lomeli, M., Hosseini, L., and Jégou, H · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Ruler: What’s the real context size of your long-context language models?
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., and Ginsburg, B · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Supervised pretraining can learn in-context reinforcement learning
Lee, J., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E · 2024
Closest in time.
Long-context llms struggle with long in-context learning
Li, T., Zhang, G., Do, Q. D., Yue, X., and Chen, W · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al · 2024
Closest in time.
XL 2 Bench: A benchmark for extremely long context understanding with long-range dependencies
Ni, X., Cai, H., Wei, X., Wang, S., Yin, D., and Li, P · 2024
Closest in time.
Transformer-based language models for reasoning in the description logic alcq
Poulis, A., Tsalapati, E., and Koubarakis, M · 2024
Closest in time.
Clongeval: A chinese benchmark for evaluating long-context large language models
Qiu, Z., Li, J., Huang, S., Zhong, W., and King, I · 2024
Closest in time.
Docfinqa: A long-context financial reasoning dataset
Reddy, V., Koncel-Kedziorski, R., Lai, V. D., and Tanner, C · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Associative recurrent memory transformer
Rodkin, I., Kuratov, Y., Bulatov, A., and Burtsev, M · 2024
Closest in time.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., et al · 2024
Closest in time.
Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k
Yuan, T., Ning, X., Zhou, D., Yang, Z., Li, S., Zhuang, M., Tan, Z., Yao, Z., Lin, D., Li, B., et al · 2024
Closest in time.