Fetching the paper…
Reading the bibliography…
Parametric language models (LMs), which are trained on vast amounts of web data, exhibit remarkable flexibility and capability.
Schwartz, R., Dodge, J., Smith, N. A., and Etzioni, O · 1907
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S. E. and Zaragoza, H · 2009
Earlier work this paper cites.
Nearest neighbor machine translation
Khandelwal, U., Fan, A., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2010
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2012
Earlier work this paper cites.
Reading Wikipedia to answer open-domain questions
Chen, D., Fisch, A., Weston, J., and Bordes, A · 2017
Earlier work this paper cites.
Billion-scale similarity search with gpus
Johnson, J., Douze, M., and Jégou, H · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Search engine guided neural machine translation
Gu, J., Wang, Y., Cho, K., and Li, V. O. K · 2018
Earlier work this paper cites.
Retrieval-based neural code generation
Hayati, S. A., Olivier, R., Avvaru, P., Yin, P., Tomasic, A., and Neubig, G · 2018
Earlier work this paper cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
ELI5: Long form question answering
Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
Lee, K., Chang, M.-W., and Toutanova, K · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in NLP
Strubell, E., Ganesh, A., and McCallum, A · 2019
Earlier work this paper cites.
Learning to infer and execute 3d shape programs
Tian, Y., Luo, A., Sun, X., Ellis, K., Freeman, W. T., Tenenbaum, J. B., and Wu, J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert
Khattab, O. and Zaharia, M · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D · 2020
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach, 2020
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Challenges in information-seeking QA: Unanswerable questions and paragraph retrieval
Asai, A. and Choi, E · 2021
Earlier work this paper cites.
One question answering model for many languages with cross-lingual dense passage retrieval
Asai, A., Yu, X., Kasai, J., and Hajishirzi, H · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
De Cao, N., Aziz, W., and Titov, I · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M · 2021
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Izacard, G. and Grave, E · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J. F. J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S. M., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N. K., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M. G., Hechtman, B. A., Weidinger, L., Gabriel, I., Isaac, W. S., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K. W., Stanway, J., Bennett, L. L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
Shuster, K., Poff, S., Chen, M., Kiela, D., and Weston, J · 2021
Earlier work this paper cites.
End-to-end training of multi-document reader and retriever for open-domain question answering
Singh, D., Reddy, S., Hamilton, W., Dyer, C., and Yogatama, D · 2021
Earlier work this paper cites.
A comprehensive survey and experimental comparison of graph-based approximate nearest neighbor search
Wang, M., Xu, X., Yue, Q., and Wang, Y · 2021
Earlier work this paper cites.
Efficient passage retrieval with hashing for open-domain question answering
Yamada, I., Asai, A., and Hajishirzi, H · 2021
Earlier work this paper cites.
Adaptive Semiparametric Language Models
Yogatama, D., de Masson d’Autume, C., and Kong, L · 2021
Earlier work this paper cites.
GPT-NeoX-20B: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S · 2022
Cited alongside, same era.
Attributed question answering: Evaluation and modeling for attributed large language models
Bohnet, B., Tran, V. Q., Verga, P., Aharoni, R., Andor, D., Soares, L. B., Eisenstein, J., Ganchev, K., Herzig, J., Hui, K., et al · 2022
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., De Las Casas, D., Guy, A., Menick, J., Ring, R., Hennigan, T., Huang, S., Maggiore, L., Jones, C., Cassirer, A., Brock, A., Paganini, M., Irving, G., Vinyals, O., Osindero, S., Simonyan, K., Rae, J., Elsen, E., and Sifre, L · 2022
Cited alongside, same era.
What does it mean for a language model to preserve privacy?
Brown, H., Lee, K., Mireshghallah, F., Shokri, R., and Tramèr, F · 2022
Cited alongside, same era.
Copy is all you need
Lan, T., Cai, D., Wang, Y., Huang, H., and Mao, X.-L · 2023
Later among the works it cites.
How to train your dragon: Diverse augmentation towards generalizable dense retrieval
Lin, S.-C., Asai, A., Li, M., Oguz, B., Lin, J., Mehdad, Y., Yih, W.-t., and Chen, X · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2023
Later among the works it cites.
Longpre, S., Yauney, G., Reif, E., Lee, K., Roberts, A., Zoph, B., Zhou, D., Wei, J., Robinson, K., Mimno, D., et al · 2023
Later among the works it cites.
Sail: Search-augmented instruction learning
Luo, H., Chuang, Y.-S., Gong, Y., Zhang, T., Kim, Y., Wu, X., Fox, D., Meng, H., and Glass, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, W., Hu, H., Chen, X., Verga, P., and Cohen, W · 2022
Cited alongside, same era.
PaLM: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Mention memory: incorporating textual knowledge into transformers through entity mention attention
de Jong, M., Zemlyanskiy, Y., FitzGerald, N., Sha, F., and Cohen, W. W · 2022
Cited alongside, same era.
You can’t pick your neighbors, or can you? when and how to rely on retrieval in the kNN-LM
Drozdov, A., Wang, S., Rahimi, R., McCallum, A., Zamani, H., and Iyyer, M · 2022
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E · 2022
Cited alongside, same era.
Lifelong pretraining: Continually adapting language models to emerging corpora
Jin, X., Zhang, D., Zhu, H., Xiao, W., Li, S.-W., Wei, X., Arnold, A., and Ren, X · 2022
Cited alongside, same era.
Overcoming catastrophic forgetting during domain adaptation of seq2seq language generation
Li, D., Chen, Z., Cho, E., Hao, J., Liu, X., Xing, F., Guo, C., and Liu, Y · 2022
Cited alongside, same era.
Z-ICL: Zero-shot in-context learning with pseudo-demonstrations
Lyu, X., Min, S., Beltagy, I., Zettlemoyer, L., and Hajishirzi, H · 2023
Later among the works it cites.
Expertqa: Expert-curated questions and attributed answers
Malaviya, C., Lee, S., Chen, S., Sieber, E., Yatskar, M., and Roth, D · 2023
Later among the works it cites.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., and Hajishirzi, H · 2023
Later among the works it cites.
Nonparametric masked language modeling
Min, S., Shi, W., Lewis, M., Chen, X., Yih, W.-t., Hajishirzi, H., and Zettlemoyer, L · 2023
Later among the works it cites.
Cross-lingual retrieval augmented prompt for low-resource languages
Nie, E., Liang, S., Schmid, H., and Schütze, H · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Fine-tuning or retrieval? comparing knowledge injection in llms
Ovadia, O., Brief, M., Mishaeli, M., and Elisha, O · 2023
Later among the works it cites.
Measuring and narrowing the compositionality gap in language models
Press, O., Zhang, M., Min, S., Schmidt, L., Smith, N., and Lewis, M · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Later among the works it cites.
Long-range language modeling with self-retrieval
Rubin, O. and Berant, J · 2023
Later among the works it cites.
Retrieval-based language models using a multi-domain datastore
Shao, R., Min, S., Zettlemoyer, L., and Koh, P. W · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
Later among the works it cites.
One embedder, any task: Instruction-finetuned text embeddings
Su, H., Shi, W., Kasai, J., Wang, Y., Hu, Y., Ostendorf, M., Yih, W.-t., Smith, N. A., Zettlemoyer, L., and Yu, T · 2023
Later among the works it cites.
Shall we pretrain autoregressive language models with retrieval? a comprehensive study
Wang, B., Ping, W., Xu, P., McAfee, L., Liu, Z., Shoeybi, M., Dong, Y., Kuchaiev, O., Li, B., Xiao, C., Anandkumar, A., and Catanzaro, B · 2023
Later among the works it cites.
Leandojo: Theorem proving with retrieval-augmented language models
Yang, K., Swope, A. M., Gu, A., Chalamala, R., Song, P., Yu, S., Godil, S., Prenger, R., and Anandkumar, A · 2023
Later among the works it cites.
ReAct: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2023
Later among the works it cites.
Automatic evaluation of attribution by large language models
Yue, X., Wang, B., Chen, Z., Zhang, K., Su, Y., and Sun, H · 2023
Later among the works it cites.
Distilling and retrieving generalizable knowledge for robot manipulation via language corrections
Zha, L., Cui, Y., Lin, L.-H., Kwon, M., Arenas, M. G., Zeng, A., Xia, F., and Sadigh, D · 2023
Later among the works it cites.
MQuAKE: Assessing knowledge editing in language models via multi-hop questions
Zhong, Z., Wu, Z., Manning, C., Potts, C., and Chen, D · 2023
Later among the works it cites.
Docprompting: Generating code by retrieving the docs
Zhou, S., Alon, U., Xu, F. F., Jiang, Z., and Neubig, G · 2023
Later among the works it cites.
Self-RAG: Learning to retrieve, generate, and critique through self-reflection
Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H · 2024
Closest in time.
Llemma: An open language model for mathematics
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2024
Closest in time.
BTR: Binary token representations for efficient retrieval augmented language models
Cao, Q., Min, S., Wang, Y., and Hajishirzi, H · 2024
Closest in time.
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P.-E., Lomeli, M., Hosseini, L., and Jégou, H · 2024
Closest in time.
Olmo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., et al · 2024
Closest in time.
The false promise of imitating proprietary language models
Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., and Song, D · 2024
Closest in time.
Rag vs fine-tuning: Pipelines, tradeoffs, and a case study on agriculture
Gupta, A., Shirgaonkar, A., Balaguer, A. d. L., Silva, B., Holstein, D., Li, D., Marsman, J., Nunes, L. O., Rouzbahman, M., Sharp, M., et al · 2024
Closest in time.
DSPy: Compiling declarative language model calls into state-of-the-art pipelines
Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., et al · 2024
Closest in time.
Talkin”bout ai generation: Copyright and the generative-ai supply chain
Lee, K., Cooper, A. F., and Grimmelmann, J · 2024
Closest in time.
RA-DIT: Retrieval-augmented dual instruction tuning
Lin, X. V., Chen, X., Chen, M., Shi, W., Lomeli, M., James, R., Rodriguez, P., Kahn, J., Szilvasy, G., Lewis, M., Zettlemoyer, L., and Yih, S · 2024
Closest in time.
SILO language models: Isolating legal risk in a nonparametric datastore
Min, S., Gururangan, S., Wallace, E., Hajishirzi, H., Smith, N. A., and Zettlemoyer, L · 2024
Closest in time.
Fine-grained hallucinations detections
Mishra, A., Asai, A., Wang, Y., Balachandran, V., Neubig, G., Tsvetkov, Y., and Hajishirzi, H · 2024
Closest in time.
Generative representational instruction tuning
Muennighoff, N., Su, H., Wang, L., Yang, N., Wei, F., Yu, T., Singh, A., and Kiela, D · 2024
Closest in time.
In-context pretraining: Language modeling beyond document boundaries
Shi, W., Min, S., Lomeli, M., Zhou, C., Li, M., Lin, V., Smith, N. A., Zettlemoyer, L., Yih, S., and Lewis, M · 2024
Closest in time.
RECOMP: Improving retrieval-augmented LMs with context compression and selective augmentation
Xu, F., Shi, W., and Choi, E · 2024
Closest in time.
Making retrieval-augmented language models robust to irrelevant context
Yoran, O., Wolfson, T., Ram, O., and Berant, J · 2024
Closest in time.