Fetching the paper…
Reading the bibliography…
The scaling of inference computation has unlocked the potential of long-context large language models (LLMs) across diverse settings.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis · 2019
Earlier work this paper cites.
Reformer: The efficient transformer
N. Kitaev, L. Kaiser, and A. Levskaya · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Rethinking attention with performers
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, et al · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang · 2020
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
X. Ho, A.-K. D. Nguyen, S. Sugawara, and A. Aizawa · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
KILT: a benchmark for knowledge intensive language tasks
F. Petroni, A. Piktus, A. Fan, P. Lewis, M. Yazdani, N. De Cao, J. Thorne, Y. Jernite, V. Karpukhin, J. Maillard, et al · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Earlier work this paper cites.
Did Aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
A. Gu, K. Goel, and C. Ré · 2021
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
G. Izacard and É. Grave · 2021
Earlier work this paper cites.
Large dual encoders are generalizable retrievers
J. Ni, C. Qu, J. Lu, Z. Dai, G. H. Ábrego, J. Ma, V. Y. Zhao, Y. Luan, K. B. Hall, M.-W. Chang, et al · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
O. Press, N. A. Smith, and M. Lewis · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, et al · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with IO-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Earlier work this paper cites.
What makes good in-context examples for GPT-3?
J. Liu, D. Shen, Y. Zhang, W. B. Dolan, L. Carin, and W. Chen · 2022
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp · 2022
Earlier work this paper cites.
MetaICL: Learning to learn in context
S. Min, M. Lewis, L. Zettlemoyer, and H. Hajishirzi · 2022
Cited alongside, same era.
Learning to retrieve prompts for in-context learning
O. Rubin, J. Herzig, and J. Berant · 2022
Cited alongside, same era.
MuSiQue: Multihop questions via single-hop question composition
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Extending context window of large language models via positional interpolation
Many-shot in-context learning
R. Agarwal, A. Singh, L. M. Zhang, B. Bohnet, L. Rosias, S. C. Chan, B. Zhang, A. Faust, and H. Larochelle · 2024
Closest in time.
xLSTM: Extended long short-term memory
M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2024
Closest in time.
In-context learning with long-context models: An in-depth exploration
A. Bertsch, M. Ivgi, U. Alon, J. Berant, M. R. Gormley, and G. Neubig · 2024
Closest in time.
HippoRAG: Neurobiologically inspired long-term memory for large language models
B. J. Gutiérrez, Y. Shu, Y. Gu, M. Yasunaga, and Y. Su · 2024
Closest in time.
LongRAG: Enhancing retrieval-augmented generation with long-context LLMs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Chen, S. Wong, L. Chen, and Y. Tian · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Cited alongside, same era.
Pre-training to learn in context
Y. Gu, L. Dong, F. Wei, and M. Huang · 2023
Cited alongside, same era.
Atlas: Few-shot learning with retrieval augmented language models
G. Izacard, P. Lewis, M. Lomeli, L. Hosseini, F. Petroni, T. Schick, J. Dwivedi-Yu, A. Joulin, S. Riedel, and E. Grave · 2023
Cited alongside, same era.
S. A. Jacobs, M. Tanaka, C. Zhang, M. Zhang, L. Song, S. Rajbhandari, and Y. He · 2023
Cited alongside, same era.
Active retrieval augmented generation
Z. Jiang, F. F. Xu, L. Gao, Z. Sun, Q. Liu, J. Dwivedi-Yu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
In-context learning with many demonstration examples
M. Li, S. Gong, J. Feng, Y. Xu, J. Zhang, Z. Wu, and L. Kong · 2023
Cited alongside, same era.
Z. Jiang, X. Ma, and W. Chen · 2024
Closest in time.
Automata-based constraints for language model decoding
T. Koo, F. Liu, and L. He · 2024
Closest in time.
BABILong: Testing the limits of LLMs with long context reasoning-in-a-haystack
Y. Kuratov, A. Bulatov, P. Anokhin, I. Rodkin, D. Sorokin, A. Sorokin, and M. Burtsev · 2024
Closest in time.
Long context rag performance of large language models
Q. Leng, J. Portes, S. Havens, M. Zaharia, and M. Carbin · 2024
Closest in time.
Long-context LLMs struggle with long in-context learning
T. Li, G. Zhang, Q. D. Do, X. Yue, and W. Chen · 2024
Closest in time.
RA-DIT: Retrieval-augmented dual instruction tuning
X. V. Lin, X. Chen, M. Chen, W. Shi, M. Lomeli, R. James, P. Rodriguez, J. Kahn, G. Szilvasy, M. Lewis, et al · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
RAPTOR: Recursive abstractive processing for tree-organized retrieval
P. Sarthi, S. Abdullah, A. Tuli, S. Khanna, A. Goldie, and C. D. Manning · 2024
Closest in time.
Scaling retrieval-based language models with a trillion-token datastore
R. Shao, J. He, A. Asai, W. Shi, T. Dettmers, S. Min, L. Zettlemoyer, and P. W. Koh · 2024
Closest in time.
REPLUG: Retrieval-augmented black-box language models
W. Shi, S. Min, M. Yasunaga, M. Seo, R. James, M. Lewis, L. Zettlemoyer, and W.-t. Yih · 2024
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Closest in time.
SOAR: improved indexing for approximate nearest neighbor search
P. Sun, D. Simcha, D. Dopson, R. Guo, and S. Kumar · 2024
Closest in time.
Large language models are latent variable models: Explaining and finding good demonstrations for in-context learning
X. Wang, W. Zhu, M. Saxon, M. Steyvers, and W. Y. Wang · 2024
Closest in time.
How faithful are RAG models? quantifying the tug-of-war between RAG and LLMs’ internal prior
K. Wu, E. Wu, and J. Zou · 2024
Closest in time.
Retrieval meets long context large language models
P. Xu, W. Ping, X. Wu, L. McAfee, C. Zhu, Z. Liu, S. Subramanian, E. Bakhturina, M. Shoeybi, and B. Catanzaro · 2024
Closest in time.
Corrective retrieval augmented generation
S.-Q. Yan, J.-C. Gu, Y. Zhu, and Z.-H. Ling · 2024
Closest in time.
Making retrieval-augmented language models robust to irrelevant context
O. Yoran, T. Wolfson, O. Ram, and J. Berant · 2024
Closest in time.
Evidence-driven retrieval augmented response generation for online misinformation
Z. Yue, H. Zeng, Y. Lu, L. Shang, Y. Zhang, and D. Wang · 2024
Closest in time.
RAFT: Adapting language model to domain specific rag
T. Zhang, S. G. Patil, N. Jain, S. Shen, M. Zaharia, I. Stoica, and J. E. Gonzalez · 2024
Closest in time.