Fetching the paper…
Reading the bibliography…
Long context inference presents challenges at the system level with increased compute and memory requirements, as well as from an accuracy perspective in being able to reason over long contexts.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
The narrativeqa reading comprehension challenge, 2017
Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2017
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Multi-news: a large-scale multi-document summarization dataset and abstractive hierarchical model, 2019
Fabbri, A. R., Li, I., She, T., Li, S., and Radev, D. R · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al · 2019
Earlier work this paper cites.
Nogueira, R. and Cho, K · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Extractive summarization of long documents by combining global and local context, 2019
Xiao, W. and Carenini, G · 2019
Earlier work this paper cites.
Power-bert: Accelerating bert inference via progressive word-vector elimination
Goyal, S., Choudhury, A. R., Raje, S., Chakaravarthy, V., Sabharwal, Y., and Verma, A · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2020
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps, 2020
Ho, X., Nguyen, A.-K. D., Sugawara, S., and Aizawa, A · 2020
Earlier work this paper cites.
Length-adaptive transformer: Train once with length drop, use anytime with search
Kim, G. and Cho, K · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers, 2021
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M · 2021
Cited alongside, same era.
Efficient attentions for long document summarization, 2021
Huang, L., Cao, S., Parulian, N., Ji, H., and Wang, L · 2021
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E · 2021
Cited alongside, same era.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Wang, H., Zhang, Z., and Han, S · 2021
Cited alongside, same era.
Qmsum: A new benchmark for query-based multi-domain meeting summarization, 2021
Zhong, M., Yin, D., Yu, T., Zaidi, A., Mutuma, M., Jha, R., Awadallah, A. H., Celikyilmaz, A., Liu, Y., Qiu, X., and Radev, D · 2021
Cited alongside, same era.
New models and developer products announced at devday 2023, Nov 2023
OpenAI · 2023
Later among the works it cites.
Memgpt: Towards llms as operating systems
Packer, C., Fang, V., Patil, S. G., Lin, K., Wooders, S., and Gonzalez, J. E · 2023
Later among the works it cites.
Open-sourcing sqleval: our framework for evaluating llm-generated sql, 2023
Ping, W. J · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., and Mann, G · 2023
Later among the works it cites.
Recomp: Improving retrieval-augmented lms with compression and selective augmentation
Xu, F., Shi, W., and Choi, E · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised dense information retrieval with contrastive learning
Gautier, I., Mathilde, C., Lucas, H., Sebastian, R., Piotr, B., Armand, J., and Edouard, G · 2022
Cited alongside, same era.
Learned token pruning for transformers
Kim, S., Shen, S., Thorsley, D., Gholami, A., Kwon, W., Hassoun, J., and Keutzer, K · 2022
Cited alongside, same era.
Musique: Multihop questions via single-hop question composition, 2022
Trivedi, H., Balasubramanian, N., Khot, T., and Sabharwal, A · 2022
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., and Wei, F · 2022
Cited alongside, same era.
Introducing claude 2.1, Nov 2023
Anthropic · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Cited alongside, same era.
Adapting language models to compress contexts
Chevalier, A., Wettig, A., Ajith, A., and Chen, D · 2023
Cited alongside, same era.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Prompt-saw: Leveraging relation-aware graphs for textual prompt compression
Ali, M. A., Li, Z., Yang, S., Cheng, K., Cao, Y., Huang, T., Hu, L., Yu, L., and Wang, D · 2024
Closest in time.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Snapkv: Llm knows what you are looking for before generation
Li, Y., Huang, Y., Yang, B., Venkitesh, B., Locatelli, A., Ye, H., Cai, T., Lewis, P., and Chen, D · 2024
Closest in time.
Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression, 2024
Pan, Z., Wu, Q., Jiang, H., Xia, M., Luo, X., Zhang, J., Lin, Q., Rühle, V., Yang, Y., Lin, C.-Y., Zhao, H. V., Qiu, L., and Zhang, D · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Boost your search with the crispy mixedbread rerank models, 2024
Shakir, A., Koenig, D., Lipp, J., and Lee, S · 2024
Closest in time.
Lloco: Learning long contexts offline
Tan, S., Li, X., Patil, S., Wu, Z., Zhang, T., Keutzer, K., Gonzalez, J. E., and Popa, R. A · 2024
Closest in time.
Introducing dbrx: A new state-of-the-art open llm, 2024
Team, T. M. R · 2024
Closest in time.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2024
Closest in time.