2024

Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models

Bordt, Sebastian, Nori, Harsha, Rodrigues, Vanessa et al.

Understand

While many have shown how Large Language Models (LLMs) can be applied to a diverse set of tasks, the critical issues of data contamination and memorization are often glossed over.

  • In this work, we address this concern for tabular data.
  • Specifically, we introduce a variety of different techniques to assess whether a language model has seen a tabular dataset during training.
  • This investigation reveals that LLMs have memorized many popular tabular datasets verbatim.

Reading the bibliography…