Fetching the paper…
Reading the bibliography…
Instruction fine-tuning attacks pose a serious threat to large language models (LLMs) by subtly embedding poisoned examples in fine-tuning datasets, leading to harmful or unintended behaviors in downstream applications.
Robust statistics—the approach based on influence functions
John Law. 1986 · 1986
Earlier work this paper cites.
Understanding black-box predictions via influence functions. In International conference on machine learning . PMLR, 1885–1894
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 2019
Earlier work this paper cites.
Influence function based data poisoning attacks to top-n recommender systems. In Proceedings of The Web Conference 2020 . 3019–3025
Minghong Fang, Neil Zhenqiang Gong, and Jia Liu. 2020 · 2020
Earlier work this paper cites.
Data collection and quality challenges for deep learning
Steven Euijong Whang and Jae-Gil Lee. 2020 · 2020
Earlier work this paper cites.
Influence Functions in Deep Learning Are Fragile. In International Conference on Learning Representations
Samyadeep Basu, Phil Pope, and Soheil Feizi. 2021 · 2021
Earlier work this paper cites.
Efficient and effective data imputation with influence functions
Xiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao, Jun Wang, and Jianwei Yin. 2021 · 2021
Earlier work this paper cites.
MathBERT: A Pre-trained Language Model for General NLP Tasks in Mathematics Education. In NeurIPS 2021 Math AI for Education Workshop
Jia Tracy Shen, Michiharu Yamashita, Ethan Prihar, Neil Heffernan, Xintao Wu, Ben Graff, and Dongwon Lee. 2021 · 2021
Earlier work this paper cites.
Interactive label cleaning with example-based explanations
Stefano Teso, Andrea Bontempelli, Fausto Giunchiglia, and Andrea Passerini. 2021 · 2021
Earlier work this paper cites.
The influence function of semiparametric estimators
Hidehiko Ichimura and Whitney K Newey. 2022 · 2022
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions. In ACL
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022 · 2022
Earlier work this paper cites.
Explainable ai: Foundations, applications, opportunities for data management research. In Proceedings of the 2022 International Conference on Management of Data . 2452–2457
Romila Pradhan, Aditya Lahiri, Sainyam Galhotra, and Babak Salimi. 2022a · 2022
Earlier work this paper cites.
Interpretable data-based explanations for fairness debugging. In Proceedings of the 2022 International Conference on Management of Data . 247–261
Romila Pradhan, Jiongli Zhu, Boris Glavic, and Babak Salimi. 2022b · 2022
Earlier work this paper cites.
Cross-loss influence functions to explain deep network representations. In International Conference on Artificial Intelligence and Statistics . PMLR, 1–17
Andrew Silva, Rohit Chopra, and Matthew Gombolay. 2022 · 2022
Cited alongside, same era.
Super-NaturalInstructions:Generalization via Declarative Instructions on 1600+ Tasks. In EMNLP
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Natural Instructions: A Repository of Language Instructions for NLP Tasks
allenai. 2023 · 2023
Cited alongside, same era.
Tracing Model Outputs to the Training Data
Anthropic. 2023 · 2023
Chatbox AI: Powerful AI Client
Benn Huang. 2024 · 2024
Later among the works it cites.
Machine unlearning in learned databases: An experimental analysis
Meghdad Kurmanji, Eleni Triantafillou, and Peter Triantafillou. 2024 · 2024
Later among the works it cites.
LLM-PBE: Assessing Data Privacy in Large Language Models
Qinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan, Rachel Xin, Junyi Hou, Xavier Yin, Zhun Wang, Dan Hendrycks, Zhangyang Wang, et al · 2024
Later among the works it cites.
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, et al · 2024
Later among the works it cites.
Chatbot Arena (formerly LMSYS): Free AI Chat to Compare & Test Best AI Chatbots
LMSYS. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Studying large language model generalization with influence functions
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Cited alongside, same era.
On the exploitability of instruction tuning
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. 2023 · 2023
Cited alongside, same era.
Poisoning Instruction-Tuned Models
Alexander Wan. 2023 · 2023
Cited alongside, same era.
Poisoning language models during instruction tuning. In International Conference on Machine Learning . PMLR, 35413–35425
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023 · 2023
Cited alongside, same era.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
Demystifying Data Management for Large Language Models. In Companion of the 2024 International Conference on Management of Data . 547–555
Xupeng Miao, Zhihao Jia, and Bin Cui. 2024 · 2024
Later among the works it cites.
Fine-Tuning Data
Nurdle. 2024 · 2024
Later among the works it cites.
Kronfluence: A PyTorch Package for Influence Function Computation
Pomonam. 2023 · 2024
Later among the works it cites.
All About Crowdsourcing Data Annotation: Leveraging the Power of the Crowd
Sapien. 2024 · 2024
Later among the works it cites.
Build an LLM-Powered Data Agent for Data Analysis
Tanay Varshney. 2024 · 2024
Later among the works it cites.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. 2024 · 2024
Later among the works it cites.
Foundation Models for Decision Making: Algorithms, Frameworks, and Applications
Sherry Yang. 2024 · 2024
Later among the works it cites.
Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) . 9304–9314
Shuo Yang and Gjergji Kasneci. 2024 · 2024
Later among the works it cites.