Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are highly sensitive to even small amounts of unsafe training data, making effective detection and filtering essential for trustworthy model development.
Calculation of signal detection theory measures
Hal Stanislaw and Natasha Todorov. 1999 · 1999
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Mukund Sundararajan, and Satyen Kale. 2020 · 2002
Earlier work this paper cites.
An introduction to signal detection and estimation
H Vincent Poor. 2013 · 2013
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2020 · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec. 2020 · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
An empirical comparison of instance attribution methods for nlp
Pouya Pezeshkpour, Sarthak Jain, Byron C. Wallace, and Sameer Singh. 2021 · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
PromDA: Prompt-based data augmentation for low-resource NLU tasks
Yufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, and Daxin Jiang. 2022 · 2022
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Earlier work this paper cites.
Analyzing the use of large language models for content moderation with chatgpt examples
Mirko Franco, Ombretta Gaggi, and Claudio E. Palazzi. 2023 · 2023
Earlier work this paper cites.
Gender bias and stereotypes in large language models
Hadas Kotek, Rikker Dockum, and David Sun. 2023 · 2023
Earlier work this paper cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. 2023 · 2023
Earlier work this paper cites.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Eloundou, Teddy Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023 · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Cited alongside, same era.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023 · 2023
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Turning generative models degenerate: The power of data poisoning attacks
Shuli Jiang, Swanand Ravindra Kadhe, Yi Zhou, Farhan Ahmed, Ling Cai, and Nathalie Baracaldo. 2024 · 2024
Later among the works it cites.
Datainf: Efficiently estimating data influence in lora-tuned llms and diffusion models
Yongchan Kwon, Eric Wu, Kevin Wu, and James Zou. 2024 · 2024
Later among the works it cites.
AI @ Meta Llama Team. 2024 · 2024
Later among the works it cites.
Less: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024 · 2024
Later among the works it cites.
Gradsafe: Detecting jailbreak prompts for llms via safety-critical gradient analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
Do large language models discriminate in hiring decisions on the basis of race, ethnicity, and gender?
Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, and Rachel Rudinger. 2024 · 2024
Cited alongside, same era.
How susceptible are large language models to ideological manipulation?
Kai Chen, Zihao He, Jun Yan, Taiwei Shi, and Kristina Lerman. 2024 · 2024
Cited alongside, same era.
Cognitive bias in decision-making with LLMs
Jessica Maria Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He. 2024 · 2024
Cited alongside, same era.
Aegis: Online adaptive ai content safety moderation with ensemble of llm experts
Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien. 2024 · 2024
Cited alongside, same era.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. 2024 · 2024
Cited alongside, same era.
What is in your safe data? identifying benign data that breaks safety
Luxi He, Mengzhou Xia, and Peter Henderson. 2024 · 2024
Cited alongside, same era.
Instructed to bias: Instruction-tuned language models exhibit emergent cognitive bias
Itay Itzhak, Gabriel Stanovsky, Nir Rosenfeld, and Yonatan Belinkov. 2024 · 2024
Cited alongside, same era.
Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Gong. 2024 · 2024
Later among the works it cites.
On the vulnerability of safety alignment in open-access LLMs
Jingwei Yi, Rui Ye, Qisi Chen, Bin Zhu, Siheng Chen, Defu Lian, Guangzhong Sun, Xing Xie, and Fangzhao Wu. 2024 · 2024
Later among the works it cites.
Shieldgemma: Generative ai content moderation based on gemma
Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, Olivia Sturman, and Oscar Wahltinez. 2024 · 2024
Later among the works it cites.
Lightweight safety guardrails using fine-tuned bert embeddings
Aaron Zheng, Mansi Rana, and Andreas Stolcke. 2024 · 2024
Later among the works it cites.
A survey of data attribution: Methods, applications, and evaluation in the era of generative ai
Junwei Deng, Yuzheng Hu, Pingbang Hu, Ting-wei Li, Shixuan Liu, Jiachen T. Wang, Dan Ley, Qirun Dai, Benhao Huang, Jin Huang, Cathy Jiao, Hoang Anh Just, Yijun Pan, Jingyan Shen, Yiwen Tu, Weiyi Wang, Xinhe Wang, Shichang Zhang, Shiyuan Zhang, Ruoxi Jia, Himabindu Lakkaraju, Hao Peng, Weijing Tang, Chenyan Xiong, Jieyu Zhao, Hanghang Tong, Han Zhao, and Jiaqi W. Ma. 2025 · 2025
Closest in time.
Date-lm: Benchmarking data attribution evaluation for large language models
Cathy Jiao, Yijun Pan, Emily Xiao, Daisy Sheng, Niket Jain, Hanzhang Zhao, Ishita Dasgupta, Jiaqi W. Ma, and Chenyan Xiong. 2025 · 2025
Closest in time.
Layer-aware representation filtering: Purifying finetuning data to preserve llm safety alignment
Hao Li, Lijun Li, Zhenghao Lu, Xianyi Wei, Rui Li, Jing Shao, and Lei Sha. 2025 · 2025
Closest in time.
Manuel Weber, Moritz Huber, Maximilian Auch, Alexander Döschl, Max-Emanuel Keller, and Peter Mandl. 2025 · 2025
Closest in time.