Fetching the paper…
Reading the bibliography…
While large language models have demonstrated impressive performance across various domains and tasks, their security issues have become increasingly severe.
The mahalanobis distance
Roy De Maesschalck, Delphine Jouan-Rimbaud, and Désiré L Massart · 2000
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Anomaly detection in sea traffic-a comparison of the gaussian mixture model and the kernel density estimator
Rikard Laxhammar, Goran Falkman, and Egils Sviestins · 2009
Earlier work this paper cites.
High-dimensional and large-scale anomaly detection using a linear one-class svm with deep learning
Sarah M Erfani, Sutharshan Rajasegarar, Shanika Karunasekera, and Christopher Leckie · 2016
Earlier work this paper cites.
Anomaly detection using random forest: A performance revisited
Rifkie Primartha and Bayu Adhi Tama · 2017
Earlier work this paper cites.
General data protection regulation
General Data Protection Regulation GDPR · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Deep autoencoding gaussian mixture model for unsupervised anomaly detection
Bo Zong, Qi Song, Martin Renqiang Min, Wei Cheng, Cristian Lumezanu, Daeki Cho, and Haifeng Chen · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
UCI Machine Learning Repository, 2020
CLINC150 · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Eternal sunshine of the spotless net: Selective forgetting in deep networks
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Privacy risks of general-purpose language models
Xudong Pan, Mi Zhang, Shouling Ji, and Min Yang · 2020
Earlier work this paper cites.
Csi: Novelty detection via contrastive learning on distributionally shifted instances
Jihoon Tack, Sangwoo Mo, Jongheon Jeong, and Jinwoo Shin · 2020
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings
T Gao, X Yao, and Danqi Chen · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Earlier work this paper cites.
Unsupervised out-of-domain detection via pre-trained transformers
Keyang Xu, Tongzheng Ren, Shikun Zhang, Yihao Feng, and Caiming Xiong · 2021
Cited alongside, same era.
Generalized out-of-distribution detection: A survey
Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu · 2021
Cited alongside, same era.
Estimating soft labels for out-of-domain intent detection
Hao Lang, Yinhe Zheng, Jian Sun, Fei Huang, Luo Si, and Yongbin Li · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Non-transferable learning: A new approach for model ownership verification and applicability authorization
Vidur: A large-scale simulation framework for llm inference
Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav Gulavani, Ramachandran Ramjee, and Alexey Tumanov · 2024
Closest in time.
Envisioning outlier exposure by large language models for out-of-distribution detection
Chentao Cao, Zhun Zhong, Zhanke Zhou, Yang Liu, Tongliang Liu, and Bo Han · 2024
Closest in time.
Do localization methods actually localize memorized data in llms? a tale of two benchmarks
Ting-Yun Chang, Jesse Thomason, and Robin Jia · 2024
Closest in time.
Learnable privacy neurons localization in language models
Ruizhe Chen, Tianxiang Hu, Yang Feng, and Zuozhu Liu · 2024
Closest in time.
Unsupervised layer-wise score aggregation for textual ood detection
Maxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie CK Cheung, Pablo Piantanida, and Pierre Colombo · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lixu Wang, Shichao Xu, Ruiqi Xu, Xiao Wang, and Qi Zhu · 2022
Cited alongside, same era.
Unlearn what you want to forget: Efficient unlearning for llms
Jiaao Chen and Diyi Yang · 2023
Cited alongside, same era.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2023
Cited alongside, same era.
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich · 2023
Cited alongside, same era.
Domain watermark: Effective and harmless dataset copyright protection is closed at hand
Junfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia, Heng Huang, Cong Liu, and Bo Li · 2023
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo · 2023
Cited alongside, same era.
On prefix-tuning for lightweight out-of-distribution detection
Yawen Ouyang, Yongchang Cao, Yuan Gao, Zhen Wu, Jianbing Zhang, and Xinyu Dai · 2023
Cited alongside, same era.
Closest in time.
Semantic-preserving adversarial example attack against bert
Chongyang Gao, Kang Gu, Soroush Vosoughi, and Shagufta Mehnaz · 2024
Closest in time.
Model editing can hurt general abilities of large language models
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng · 2024
Closest in time.
Model editing at scale leads to gradual and catastrophic forgetting
Akshat Gupta, Anurag Rao, and Gopala Anumanchipalli · 2024
Closest in time.
Jogging the memory of unlearned models through targeted relearning attacks
Shengyuan Hu, Yiwei Fu, Steven Wu, and Virginia Smith · 2024
Closest in time.
Ravel: Evaluating interpretability methods on disentangling language model representations
Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, and Atticus Geiger · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
Soul: Unlocking the power of second-order optimization for llm unlearning
Jinghan Jia, Yihua Zhang, Yimeng Zhang, Jiancheng Liu, Bharat Runwal, James Diffenderfer, Bhavya Kailkhura, and Sijia Liu · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al · 2024
Closest in time.
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Hang Li, Kush R Varshney, et al · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter · 2024
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard · 2024
Closest in time.
In-context unlearning: Language models as few shot unlearners
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju · 2024
Closest in time.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
Guardrail baselines for unlearning in llms
Pratiksha Thaker, Yash Maurya, and Virginia Smith · 2024
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2024
Closest in time.
Negative preference optimization: From catastrophic collapse to effective unlearning
Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei · 2024
Closest in time.
Can we trust embodied agents? exploring backdoor attacks against embodied llm-based decision-making systems
Ruochen Jiao, Shaoyuan Xie, Justin Yue, TAKAMI SATO, Lixu Wang, Yixuan Wang, Qi Alfred Chen, and Qi Zhu · 2025
Closest in time.