Fetching the paper…
Reading the bibliography…
Pre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications.
Basic principles of ROC analysis. In Seminars in nuclear medicine
Charles E Metz. 1978 · 1978
Earlier work this paper cites.
Neural smithing: supervised learning in feedforward artificial neural networks
Russell Reed and Robert J MarksII. 1999 · 1999
Earlier work this paper cites.
An introduction to ROC analysis
Tom Fawcett. 2006 · 2006
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
On the definiteness of earth mover’s distance and its relation to set intersection
Andrew Gardner, Christian A Duncan, Jinko Kanno, and Rastko R Selmic. 2017 · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP)
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting. In 2018 IEEE 31st computer security foundations symposium (CSF)
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018 · 2018
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Estimating the success of re-identifications in incomplete datasets using generative models
Luc Rocher, Julien M Hendrickx, and Yves-Alexandre De Montjoye. 2019 · 2019
Cited alongside, same era.
On losses for modern language models
Stéphane Aroca-Ouellette and Frank Rudzicz. 2020 · 2020
Cited alongside, same era.
Information leakage in embedding models. In Proceedings of the 2020 ACM SIGSAC conference on computer and communications security
Congzheng Song and Ananth Raghunathan. 2020 · 2020
Cited alongside, same era.
Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21)
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Cited alongside, same era.
Enhanced membership inference attacks against machine learning models. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security
Jiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler, and Reza Shokri. 2022 · 2022
Later among the works it cites.
Z-Library
2023 · 2023
Later among the works it cites.
Speak, memory: An archaeology of books known to chatgpt/gpt-4
Kent K Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023 · 2023
Later among the works it cites.
goodreads-Popular quotes
Goodreads. 2023 · 2023
Later among the works it cites.
ChatGPT reaches 100 million users two months after launch
Guardian. 2023 · 2023
Later among the works it cites.
Do language models plagiarize?. In Proceedings of the ACM Web Conference 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huseyin A Inan, Osman Ramadan, Lukas Wutschitz, Daniel Jones, Victor Rühle, James Withers, and Robert Sim. 2021 · 2021
Cited alongside, same era.
Diffusion earth mover’s distance and distribution embeddings. In International Conference on Machine Learning
Alexander Y Tong, Guillaume Huguet, Amine Natik, Kincaid MacDonald, Manik Kuchroo, Ronald Coifman, Guy Wolf, and Smita Krishnaswamy. 2021 · 2021
Cited alongside, same era.
Are Clinical BERT Models Privacy Preserving? The Difficulty of Extracting Patient-Condition Associations.. In HUMAN@ AAAI Fall Symposium
Thomas Vakili and Hercules Dalianis. 2021 · 2021
Cited alongside, same era.
Approximating the Earth Mover’s Distance between sets of geometric objects
Marc van Kreveld, Frank Staals, Amir Vaxman, and Jordi Vermeulen. 2021 · 2021
Cited alongside, same era.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Cited alongside, same era.
Quantifying privacy risks of masked language models using membership inference attacks
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri. 2022 · 2022
Cited alongside, same era.
Jooyoung Lee, Thai Le, Jinghui Chen, and Dongwon Lee. 2023 · 2023
Later among the works it cites.
Analyzing leakage of personally identifiable information in language models
Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-Béguelin. 2023 · 2023
Later among the works it cites.
Getty Images lawsuit says Stability AI misused photos to train AI
REUTERS. 2023-02-06 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019 · 2023
Later among the works it cites.