Fetching the paper…
Reading the bibliography…
Pretraining Large Language Models (LLMs) on large corpora of textual data is now a standard paradigm.
Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts
Dernoncourt, F. and Lee, J. Y · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail
Feldman, V · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Privacy risks of general-purpose language models
Pan, X., Zhang, M., Ji, S., and Yang, M · 2020
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2022
Cited alongside, same era.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2022
Cited alongside, same era.
Deduplicating training data mitigates privacy risks in language models
Kandpal, N., Wallace, E., and Raffel, C · 2022
Cited alongside, same era.
Internet-augmented language models through few-shot prompting for open-domain question answering
Lazaridou, A., Gribovskaya, E., Stokowiec, W., and Grigorev, N · 2022
Cited alongside, same era.
Atlas: Few-shot learning with retrieval augmented language models
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2023
Later among the works it cites.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Later among the works it cites.
Freshllms: Refreshing large language models with search engine augmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards understanding grokking: An effective theory of representation learning
Liu, Z., Kitouni, O., Nolte, N. S., Michaud, E., Tegmark, M., and Williams, M · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Cited alongside, same era.
Memorisation versus generalisation in pre-trained language models
Tänzer, M., Ruder, S., and Rei, M · 2022
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Prompt engineering for claude’s long context window
Anthropic · 2023
Cited alongside, same era.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H · 2023
Cited alongside, same era.
Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., et al · 2023
Later among the works it cites.
Instructretro: Instruction tuning post retrieval-augmented pretraining
Wang, B., Ping, W., McAfee, L., Xu, P., Li, B., Shoeybi, M., and Catanzaro, B · 2023
Later among the works it cites.
System 2 attention (is something you might need too)
Weston, J. and Sukhbaatar, S · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., et al · 2023
Later among the works it cites.
Retrieval meets long context large language models
Xu, P., Ping, W., Wu, X., McAfee, L., Zhu, C., Liu, Z., Subramanian, S., Bakhturina, E., Shoeybi, M., and Catanzaro, B · 2023
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Later among the works it cites.
Chatqa: Building gpt-4 level conversational qa models
Liu, Z., Ping, W., Roy, R., Xu, P., Shoeybi, M., and Catanzaro, B · 2024
Closest in time.