Fetching the paper…
Reading the bibliography…
Substantial research works have shown that deep models, e.g., pre-trained models, on the large corpus can learn universal language representations, which are beneficial for downstream NLP tasks.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2004
Earlier work this paper cites.
Natural backdoor attack on text data
Lichao Sun. 2020 · 2006
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018 · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2019 · 2019
Earlier work this paper cites.
Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning
Milad Nasr, Reza Shokri, and Amir Houmansadr. 2019 · 2019
Earlier work this paper cites.
Ml-leaks: Model and data independent membership inference attacks and defenses on machine learning models
Ahmed Salem, Yang Zhang, Mathias Humbert, Mario Fritz, and Michael Backes. 2019 · 2019
Earlier work this paper cites.
Auditing data provenance in text-generation models
Congzheng Song and Vitaly Shmatikov. 2019 · 2019
Earlier work this paper cites.
Segmentations-leak: Membership inference attacks and defenses in semantic image segmentation
Yang He, Shadi Rahimian, Bernt Schiele, and Mario Fritz. 2020 · 2020
Earlier work this paper cites.
Membership inference attacks on sequence-to-sequence models: Is my data in your machine translation system?
Sorami Hisamoto, Matt Post, and Kevin Duh. 2020 · 2020
Earlier work this paper cites.
Information leakage in embedding models
Congzheng Song and Ananth Raghunathan. 2020 · 2020
Earlier work this paper cites.
Badnl: Backdoor attacks against nlp models
Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, and Yang Zhang. 2021a · 2021
Earlier work this paper cites.
Triggerless backdoor attack for nlp tasks with clean labels
Leilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Shangwei Guo, and Chun Fan. 2021 · 2021
Cited alongside, same era.
Mlcapsule: Guarded offline deployment of machine learning as a service
Lucjan Hanzlik, Yang Zhang, Kathrin Grosse, Ahmed Salem, Maximilian Augustin, Michael Backes, and Mario Fritz. 2021 · 2021
Cited alongside, same era.
Membership inference attacks on machine learning: A survey
Hongsheng Hu, Zoran Salcic, Lichao Sun, Gillian Dobbie, Philip S Yu, and Xuyun Zhang. 2021 · 2021
Cited alongside, same era.
Membership inference attack susceptibility of clinical language models
Abhyuday Jagannatha, Bhanu Pratap Singh Rawat, and Hong Yu. 2021 · 2021
Cited alongside, same era.
Deep learning–based text classification: a comprehensive review
Shervin Minaee, Nal Kalchbrenner, Erik Cambria, Narjes Nikzad, Meysam Chenaghlu, and Jianfeng Gao. 2021 · 2021
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Textual backdoor attacks with iterative trigger injection
Jun Yan, Vansh Gupta, and Xiang Ren. 2022 · 2022
Later among the works it cites.
Yihan Cao, Siyu Li, Yixin Liu, Zhiling Yan, Yutong Dai, Philip S Yu, and Lichao Sun. 2023 · 2023
Closest in time.
Elon musk threatens to sue microsoft over using twitter data for its a.i
Kifleswing. 2023 · 2023
Closest in time.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Privacy regularization: Joint privacy-utility optimization in language models
Fatemehsadat Mireshghallah, Huseyin A Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, and Robert Sim. 2021 · 2021
Cited alongside, same era.
Defending medical image diagnostics against privacy attacks using generative methods: Application to retinal diagnostics
William Paul, Yinzhi Cao, Miaomiao Zhang, and Phil Burlina. 2021 · 2021
Cited alongside, same era.
On the difficulty of membership inference attacks
Shahbaz Rezaei and Xin Liu. 2021 · 2021
Cited alongside, same era.
Membership inference attacks against nlp classification models
Virat Shejwalkar, Huseyin A Inan, Amir Houmansadr, and Robert Sim. 2021 · 2021
Cited alongside, same era.
Systematic evaluation of privacy risks of machine learning models
Liwei Song and Prateek Mittal. 2021 · 2021
Cited alongside, same era.
Leveraging adversarial examples to quantify membership information leakage
Ganesh Del Grosso, Hamid Jalalzai, Georg Pichler, Catuscia Palamidessi, and Pablo Piantanida. 2022 · 2022
Cited alongside, same era.
Privacy leakage in text classification: A data extraction approach
Adel Elmahdy, Huseyin A Inan, and Robert Sim. 2022 · 2022
Cited alongside, same era.
Closest in time.
Black-box dataset ownership verification via backdoor watermarking
Yiming Li, Mingyan Zhu, Xue Yang, Yong Jiang, Tao Wei, and Shu-Tao Xia. 2023 · 2023
Closest in time.
Augmented language models: a survey
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al. 2023 · 2023
Closest in time.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Closest in time.
Backdoor cleansing with unlabeled data
Lu Pang, Tao Sun, Haibin Ling, and Chao Chen. 2023 · 2023
Closest in time.
Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt
Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun. 2023 · 2023
Closest in time.
Reddit will start charging for api access to rebuff llms
Brandon Vigliarolo. 2023 · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al. 2023 · 2023
Closest in time.
Encodermi: Membership inference against pre-trained encoders in contrastive learning
Hongbin Liu, Jinyuan Jia, Wenjie Qu, and Neil Zhenqiang Gong. 2021 · 2095
Closest in time.