Fetching the paper…
Reading the bibliography…
Language Models (LMs) typically adhere to a "pre-training and fine-tuning" paradigm, where a universal pre-trained model can be fine-tuned to cater to various specialized domains.
N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramèr, “Membership inference attacks from first principles,” in 2022 IEEE Symposium on Security and Privacy (SP) , 2022, pp. 1897–1914
1914
Earlier work this paper cites.
C. A. C. Choo, F. Tramèr, N. Carlini, and N. Papernot, “Label-Only Membership Inference Attacks,” in International Conference on Machine Learning (ICML) . PMLR, 2021, pp. 1964–1974
1974
Earlier work this paper cites.
A. Krogh and J. Hertz, “A simple weight decay can improve generalization,” Advances in neural information processing systems , vol. 4, 1991
1991
Earlier work this paper cites.
C. Dwork, F. McSherry, K. Nissim, and A. Smith, “Calibrating noise to sensitivity in private data analysis,” in Theory of Cryptography , S. Halevi and T. Rabin, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2006, pp. 265–284
2006
Earlier work this paper cites.
2009
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research , vol. 15, no. 56, pp. 1929–1958, 2014. [Online]. Available: http://jmlr.org/papers/v15/srivastava14a.html
2014
Earlier work this paper cites.
X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” in Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2015. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2015/file/250cf8b51c773f3f8dc8b4be867a9a02-Paper.pdf
2015
Earlier work this paper cites.
R. Shokri, M. Stronati, C. Song, and V. Shmatikov, “Membership Inference Attacks Against Machine Learning Models,” in IEEE Symposium on Security and Privacy (S&P) . IEEE, 2017, pp. 3–18
2017
Earlier work this paper cites.
S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha, “Privacy risk in machine learning: Analyzing the connection to overfitting,” in 2018 IEEE 31st Computer Security Foundations Symposium (CSF) , 2018, pp. 268–282
2018
Earlier work this paper cites.
M. Nasr, R. Shokri, and A. Houmansadr, “Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated Learning,” in IEEE Symposium on Security and Privacy (S&P) . IEEE, 2019, pp. 1021–1035
2019
Earlier work this paper cites.
L. Song, R. Shokri, and P. Mittal, “Privacy risks of securing machine learning models against adversarial examples,” in Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security , ser. CCS ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 241–257. [Online]. Available: https://doi.org/10.1145/3319535.3354211
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=Bkg6RiCqY7
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
K. Leino and M. Fredrikson, “Stolen memories: Leveraging model memorization for calibrated White-Box membership inference,” in 29th USENIX Security Symposium (USENIX Security 20) . USENIX Association, Aug. 2020, pp. 1605–1622. [Online]. Available: https://www.usenix.org/conference/usenixsecurity20/presentation/leino
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, no. 140, pp. 1–67, 2020. [Online]. Available: http://jmlr.org/papers/v21/20-074.html
2020
Earlier work this paper cites.
OpenAI, “Chat markup language,” https://github.com/openai/openai-python/blob/284c1799070c723c6a553337134148a7ab088dd8/chatml.md
2020
Earlier work this paper cites.
Y. Kaya, S. Hong, and T. Dumitras, “On the effectiveness of regularization against membership inference attacks,” 2020
2020
Earlier work this paper cites.
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching et al. , “Trl: Transformer reinforcement learning,” https://github.com/huggingface/trl
2020
Earlier work this paper cites.
2021
Cited alongside, same era.
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson et al. , “Extracting training data from large language models,” in 30th USENIX Security Symposium (USENIX Security 21) , 2021, pp. 2633–2650
2021
Cited alongside, same era.
Z. Li and Y. Zhang, “Membership Leakage in Label-Only Exposures,” in ACM SIGSAC Conference on Computer and Communications Security (CCS) . ACM, 2021, pp. 880–895
2021
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , M.-F. Moens, X. Huang, L. Specia, and S. W.-t. Yih, Eds. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 3045–3059. [Online]. Available: https://aclanthology.org/2021.emnlp-main.243
M. Li, J. Wang, J. Wang, and S. Neel, “MoPe: Model perturbation based privacy attacks on language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 13 647–13 660. [Online]. Available: https://aclanthology.org/2023.emnlp-main.842
2023
Later among the works it cites.
W. Shi, A. Ajith, M. Xia, Y. Huang, D. Liu, T. Blevins, D. Chen, and L. Zettlemoyer, “Detecting pretraining data from large language models,” 2023
2023
Later among the works it cites.
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff et al. , “Pythia: A suite for analyzing large language models across training and scaling,” in International Conference on Machine Learning . PMLR, 2023, pp. 2397–2430
2023
Later among the works it cites.
OpenAssistant, “Openassistant top-1 conversation threads,” 2023. [Online]. Available: https://huggingface.co/datasets/OpenAssistant/oasst_top1_2023-08-25
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
OpenAI, “Chatgpt,” https://openai.com/index/chatgpt/
2022
Cited alongside, same era.
S. Mangrulkar, S. Gugger, L. Debut, Y. Belkada, S. Paul, and B. Bossan, “Peft: State-of-the-art parameter-efficient fine-tuning methods,” https://github.com/huggingface/peft
2022
Cited alongside, same era.
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
2022
Cited alongside, same era.
F. Mireshghallah, K. Goyal, A. Uniyal, T. Berg-Kirkpatrick, and R. Shokri, “Quantifying privacy risks of masked language models using membership inference attacks,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 8332–8347. [Online]. Available: https://aclanthology.org/2022.emnlp-main.570
2022
Cited alongside, same era.
2022
Cited alongside, same era.
D. Yu, S. Naik, A. Backurs, S. Gopi, H. A. Inan, G. Kamath, J. Kulkarni, Y. T. Lee, A. Manoel, L. Wutschitz, S. Yekhanin, and H. Zhang, “Differentially private fine-tuning of language models,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=Q42f0dfjECO
2022
Cited alongside, same era.
L. Wutschitz, H. A. Inan, and A. Manoel, “dp-transformers: Training transformer models with differential privacy,” https://www.microsoft.com/en-us/research/project/dp-transformers
2022
Cited alongside, same era.
H. Liu, D. Tam, M. Muqeeth, J. Mohta et al. , “Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave et al. , Eds., vol. 35. Curran Associates, Inc., 2022, pp. 1950–1965. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/0cde695b83bd186c1fd456302888454c-Paper-Conference.pdf
2022
Cited alongside, same era.
2023
Later among the works it cites.
J. Belveze, “Tl;dr news dataset,” 2023. [Online]. Available: https://huggingface.co/datasets/JulesBelveze/tldr_news
2023
Later among the works it cites.
D. Cheng, S. Huang, and F. Wei, “Adapting large language models via reading comprehension,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=y886UXPEZ0
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
R. Liu, T. Wang, Y. Cao, and L. Xiong, “Precurious: How innocent pre-trained language models turn into privacy traps,” 2024
2024
Later among the works it cites.
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi, S. ES, S. Suri, D. Glushkov, A. Dantuluri, A. Maguire, C. Schuhmann, H. Nguyen, and A. Mattick, “Openassistant conversations - democratizing large language model alignment,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2024
2024
Later among the works it cites.
W. Fu, H. Wang, C. Gao, G. Liu, Y. Li, and T. Jiang, “Membership inference attacks against fine-tuned large language models via self-prompt calibration,” in Advances in Neural Information Processing Systems , A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 2024, pp. 134 981–135 010. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2024/file/f36ad694188bb4c4bbbd61e2038e069e-Paper-Conference.pdf
2024
Later among the works it cites.
2024
Later among the works it cites.
J. G. Wang, J. Wang, M. Li, and S. Neel, “Pandora’s white-box: Increased training data leakage in open llms,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
OpenAI, “Openai api reference,” 2024, accessed: 2024-09-05. [Online]. Available: https://platform.openai.com/docs/api-reference/chat
2024
Later among the works it cites.
H. Face, “Hugging face api inference documentation,” 2024, accessed: 2024-09-05. [Online]. Available: https://huggingface.co/docs/api-inference/index
2024
Later among the works it cites.