Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Later among the works it cites.
{ \{ Updates-Leak } \} : Data Set Inference and Reconstruction Attacks in Online Learning. In 29th USENIX Security Symposium (USENIX Security 20) . 1291–1308
Ahmed Salem, Apratim Bhattacharya, Michael Backes, Mario Fritz, and Yang Zhang. 2020 · 2020
Later among the works it cites.
Songmass: Automatic song writing with pre-training and alignment constraint
Original
Zhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren, Wei Ye, Shikun Zhang, and Tao Qin. 2020 · 2020
Later among the works it cites.
Authorship attribution for neural text generation. In Conf. on Empirical Methods in Natural Language Processing (EMNLP)
Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. 2020 · 2020
Later among the works it cites.
Cord-19: The covid-19 open research dataset
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, et al · 2020
Later among the works it cites.
Attacking neural text detectors
Original
Max Wolff and Stuart Wolff. 2020 · 2020
Later among the works it cites.
Analyzing information leakage of updates to natural language models. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security . 363–375
Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Victor Rühle, Andrew Paverd, Olga Ohrimenko, Boris Köpf, and Marc Brockschmidt. 2020 · 2020
Later among the works it cites.
Detecting Fake News using Machine Learning: A Systematic Literature Review
Original
Alim Al Ayub Ahmed, Ayman Aljabouh, Praveen Kumar Donepudi, and Myung Suh Choi. 2021 · 2021
Later among the works it cites.
Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Later among the works it cites.
Artificial text detection via examining the topology of attention maps
Original
Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, and Evgeny Burnaev. 2021 · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Original
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021 · 2021
Later among the works it cites.
Investigating Memorization of Conspiracy Theories in Text Generation
Original
Sharon Levy, Michael Saxon, and William Yang Wang. 2021 · 2021
Later among the works it cites.
How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN
Original
R Thomas McCoy, Paul Smolensky, Tal Linzen, Jianfeng Gao, and Asli Celikyilmaz. 2021 · 2021
Later among the works it cites.
Privacy regularization: Joint privacy-utility optimization in language models
Original
Fatemehsadat Mireshghallah, Huseyin A Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, and Robert Sim. 2021 · 2021
Later among the works it cites.
TuringBench: A Benchmark Environment for Turing Test in the Age of Neural Text Generation. In Findings of Conf. on Empirical Methods in Natural Language Processing (EMNLP-Findings)
Adaku Uchendu, Zeyu Ma, Thai Le, Rui Zhang, and Dongwon Lee. 2021 · 2021
Later among the works it cites.
Counterfactual Memorization in Neural Language Models
Original
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2021 · 2021
Later among the works it cites.
What Does it Mean for a Language Model to Preserve Privacy?
Original
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr. 2022 · 2022
Closest in time.
Quantifying Memorization Across Neural Language Models
Original
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Closest in time.
Deduplicating training data mitigates privacy risks in language models
Original
Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022 · 2022
Closest in time.