Fetching the paper…
Reading the bibliography…
This work quantifies the risk of training data leakage from LLMs (Large Language Models) using sequence-level probabilities.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail
Feldman, V · 2020
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2022
Earlier work this paper cites.
Are large pre-trained language models leaking your personal information?
Huang, J., Shao, H., and Chang, K. C.-C · 2022
Earlier work this paper cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
Tirumala, K., Markosyan, A., Zettlemoyer, L., and Aghajanyan, A · 2022
Earlier work this paper cites.
Chat Mistral
AI, M · 2023
Earlier work this paper cites.
Sok: Memorization in general-purpose large language models
Hartmann, V., Suri, A., Bindschaedler, V., Evans, D., Tople, S., and West, R · 2023
Cited alongside, same era.
Preventing generation of verbatim memorization in language models gives a false sense of privacy
Ippolito, D., Tramer, F., Nasr, M., Zhang, C., Jagielski, M., Lee, K., Choquette Choo, C., and Carlini, N · 2023
Cited alongside, same era.
Analyzing leakage of personally identifiable information in language models
Lukas, N., Salem, A., Sim, R., Tople, S., Wutschitz, L., and Zanella-Béguelin, S · 2023
Cited alongside, same era.
Can neural network memorization be localized?
Maini, P., Mozer, M. C., Sedghi, H., Lipton, Z. C., Kolter, J. Z., and Zhang, C · 2023
Cited alongside, same era.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L · 2023
Cited alongside, same era.
Emergent and predictable memorization in large language models
Biderman, S., Prashanth, U., Sutawika, L., Schoelkopf, H., Anthony, Q., Purohit, S., and Raff, E · 2024
Closest in time.
Do membership inference attacks work on large language models?
Duan, M., Suri, A., Mireshghallah, N., Min, S., Shi, W., Zettlemoyer, L., Tsvetkov, Y., Choi, Y., Evans, D., and Hajishirzi, H · 2024
Closest in time.
Measuring memorization through probabilistic discoverable extraction
Hayes, J., Swanberg, M., Chaudhari, H., Yona, I., and Shumailov, I · 2024
Closest in time.
Alpaca against vicuna: Using llms to uncover memorization of llms
Kassem, A. M., Mahmoud, O., Mireshghallah, N., Kim, H., Tsvetkov, Y., Choi, Y., Saad, S., and Rana, S · 2024
Closest in time.
Propile: Probing privacy leakage in large language models
Kim, S., Yun, S., Lee, H., Gubri, M., Yoon, S., and Oh, S. J · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu, W., Pang, T., Liu, Q., Du, C., Kang, B., Huang, Y., Lin, M., and Yan, S · 2023
Cited alongside, same era.
Closest in time.
Scaling laws for fact memorization of large language models
Lu, X., Li, X., Cheng, Q., Ding, K., Huang, X., and Qiu, X · 2024
Closest in time.