Fetching the paper…
Reading the bibliography…
Large language models (LLMs) encapsulate a vast amount of factual information within their pre-trained weights, as evidenced by their ability to answer diverse questions across different domains.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y · 2013
Earlier work this paper cites.
Wikiextractor
Attardi, G · 2015
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Johnson, J., Douze, M., and Jégou, H · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Recall and learn: Fine-tuning deep pretrained language models with less forgetting
Chen, S., Hou, Y., Cui, Y., Che, W., Liu, T., and Yu, X · 2020
Earlier work this paper cites.
Lauscher, A., Majewska, O., Ribeiro, L. F., Gurevych, I., Rozanov, N., and Glavaš, G · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
K-bert: Enabling language representation with knowledge graph
Liu, W., Zhou, P., Zhao, Z., Wang, Z., Ju, Q., Deng, H., and Wang, P · 2020
Earlier work this paper cites.
K-adapter: Infusing knowledge into pre-trained models with adapters
Wang, R., Tang, D., Duan, N., Wei, Z., Huang, X., Cao, G., Jiang, D., Zhou, M., et al · 2020
Earlier work this paper cites.
Meta-learning via language model in-context tuning
Chen, Y., Zhong, R., Zha, S., Karypis, G., and He, H · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2021
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Text data augmentation for deep learning
Shorten, C., Khoshgoftaar, T. M., and Furht, B · 2021
Cited alongside, same era.
Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction
Chen, X., Zhang, N., Xie, X., Deng, S., Yao, Y., Tan, C., Huang, F., Si, L., and Chen, H · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Huang, Q., Tao, M., An, Z., Zhang, C., Jiang, C., Chen, Z., Wu, Z., and Feng, Y · 2023
Closest in time.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Closest in time.
Large language models struggle to learn long-tail knowledge
Kandpal, N., Deng, H., Roberts, A., Wallace, E., and Raffel, C · 2023
Closest in time.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can language models learn from explanations in context?
Lampinen, A. K., Dasgupta, I., Chan, S. C., Matthewson, K., Tessler, M. H., Creswell, A., McClelland, J. L., Wang, J. X., and Hill, F · 2022
Cited alongside, same era.
Text and code embeddings by contrastive pre-training
Neelakantan, A., Xu, T., Puri, R., Radford, A., Han, J. M., Tworek, J., Yuan, Q., Tezak, N. A., Kim, J. W., Hallacy, C., Heidecke, J., Shyam, P., Power, B., Nekoul, T. E., Sastry, G., Krueger, G., Schnurr, D. P., Such, F. P., Hsu, K. S.-K., Thompson, M., Khan, T., Sherbakov, T., Jang, J., Welinder, P., and Weng, L · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Systematic review of effect of data augmentation using paraphrasing on named entity recognition
Sharma, S., Joshi, A., Mukhija, N., Zhao, Y., Bhathena, H., Singh, P., Santhanam, S., and Biswas, P · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Cited alongside, same era.
Memorization without overfitting: Analyzing the training dynamics of large language models
Tirumala, K., Markosyan, A. H., Zettlemoyer, L., and Aghajanyan, A · 2022
Cited alongside, same era.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al · 2022
Cited alongside, same era.
Mitra, A., Del Corro, L., Mahajan, S., Codas, A., Simoes, C., Agrawal, S., Chen, X., Razdaibiedina, A., Jones, E., Aggarwal, K., et al · 2023
Closest in time.
A brief report on lawgpt 1.0: A virtual legal assistant based on gpt-3
Nguyen, H.-T · 2023
Closest in time.
Capabilities of gpt-4 on medical challenge problems
Nori, H., King, N., McKinney, S. M., Carignan, D., and Horvitz, E · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Closest in time.
Can chatgpt replace traditional kbqa models? an in-depth analysis of the question answering performance of the gpt llm family
Tan, Y., Min, D., Li, Y., Li, W., Hu, N., Chen, Y., and Qi, G · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Zephyr: Direct distillation of lm alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., et al · 2023
Closest in time.
Survey on factuality in large language models: Knowledge, retrieval and domain-specificity
Wang, C., Liu, X., Yue, Y., Tang, X., Zhang, T., Jiayang, C., Yao, Y., Gao, W., Hu, X., Qi, Z., et al · 2023
Closest in time.
C-pack: Packaged resources to advance general chinese embedding, 2023
Xiao, S., Liu, Z., Zhang, P., and Muennighoff, N · 2023
Closest in time.
Fingpt: Open-source financial large language models
Yang, H., Liu, X.-Y., and Wang, C. D · 2023
Closest in time.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Closest in time.