2022

Large Language Models Struggle to Learn Long-Tail Knowledge

Kandpal, Nikhil, Deng, Haikang, Roberts, Adam et al.

Understand

The Internet contains a wealth of knowledge -- from the birthdays of historical figures to tutorials on how to code -- all of which may be learned by language models.

  • However, while certain pieces of information are ubiquitous on the web, others appear extremely rarely.
  • In this paper, we study the relationship between the knowledge memorized by large language models and the information in pre-training datasets scraped from the web.
  • In particular, we show that a language model's ability to answer a fact-based question relates to how many documents associated with that question were seen during pre-training.

Reading the bibliography…