Fetching the paper…
Reading the bibliography…
This paper investigates using knowledge editing techniques to detoxify Large Language Models (LLMs).
Principal component analysis
Svante Wold, Kim Esbensen, and Paul Geladi. 1987 · 1987
Earlier work this paper cites.
Intraoperative neurophysiological monitoring
Jaime R Lopez. 1996 · 1996
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Back to the future: Unsupervised backprop-based decoding for counterfactual and abductive commonsense reasoning
Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras, Antoine Bosselut, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022 · 2022
Earlier work this paper cites.
Aging with GRACE: lifelong model editing with discrete key-value adaptors
Thomas Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim, and Marzyeh Ghassemi. 2022 · 2022
Earlier work this paper cites.
Probing classifiers are unreliable for concept removal and detection
Abhinav Kumar, Chenhao Tan, and Amit Sharma. 2022 · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Fast model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2022a · 2022
Earlier work this paper cites.
Memory-based model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, and Chelsea Finn. 2022b · 2022
Earlier work this paper cites.
Dune: Dataset for unified editing
Afra Feyza Akyürek, Eric Pan, Garry Kuwanto, and Derry Wijaya. 2023 · 2023
Earlier work this paper cites.
LEACE: perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023 · 2023
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned LLM
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023 · 2023
Earlier work this paper cites.
Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2023 · 2023
Earlier work this paper cites.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2023 · 2023
Earlier work this paper cites.
Opencompass: A universal evaluation platform for foundation models
OpenCompass Contributors. 2023 · 2023
Earlier work this paper cites.
Recent advances towards safe, responsible, and moral dialogue systems: A survey
Jiawen Deng, Hao Sun, Zhexin Zhang, Jiale Cheng, and Minlie Huang. 2023 · 2023
Earlier work this paper cites.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Earlier work this paper cites.
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang. 2023 · 2023
Earlier work this paper cites.
Zhangyin Feng, Weitao Ma, Weijiang Yu, Lei Huang, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023 · 2023
Earlier work this paper cites.
Editing commonsense knowledge in GPT
Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao, Xiang Lorraine Li, Sarah Wiegreffe, and Niket Tandon. 2023 · 2023
Earlier work this paper cites.
Detoxifying text with marco: Controllable revision with experts and anti-experts
Skyler Hallinan, Alisa Liu, Yejin Choi, and Maarten Sap. 2023 · 2023
Earlier work this paper cites.
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023 · 2023
Cited alongside, same era.
Xinshuo Hu, Dongfang Li, Zihao Zheng, Zhenyu Liu, Baotian Hu, and Min Zhang. 2023 · 2023
Cited alongside, same era.
Transformer-patcher: One mistake worth one neuron
Zeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou, Wenge Rong, and Zhang Xiong. 2023c · 2023
Cited alongside, same era.
Knowledge sanitization of large language models
Yoichi Ishibashi and Hidetoshi Shimodaira. 2023 · 2023
Cited alongside, same era.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023 · 2023
Later among the works it cites.
Language anisotropic cross-lingual model editing
Yang Xu, Yutai Hou, Wanxiang Che, and Min Zhang. 2023 · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang. 2023c · 2023
Later among the works it cites.
Unpacking the ethical value alignment in big models
Xiaoyuan Yi, Jing Yao, Xiting Wang, and Xing Xie. 2023 · 2023
Later among the works it cites.
Mil-decoding: Detoxifying language models at token-level via multiple instance learning
Xu Zhang and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Self-detoxifying language models via toxification reversal
Chak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang, and Wenjie Li. 2023 · 2023
Cited alongside, same era.
Evaluating dependencies in fact editing for language models: Specificity and implication awareness
Zichao Li, Ines Arous, Siva Reddy, and Jackie Chi Kit Cheung. 2023c · 2023
Cited alongside, same era.
Jailbreaking chatgpt via prompt engineering: An empirical study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023 · 2023
Cited alongside, same era.
A survey on knowledge editing of neural networks
Vittorio Mazzia, Alessandro Pedrani, Andrea Caciolai, Kay Rottmann, and Davide Bernardi. 2023 · 2023
Cited alongside, same era.
Using in-context learning to improve dialogue safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tur. 2023 · 2023
Cited alongside, same era.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau. 2023 · 2023
Cited alongside, same era.
Testing language model agents safely in the wild
Silen Naihin, David Atkinson, Marc Green, Merwane Hamadi, Craig Swift, Douglas Schonholtz, Adam Tauman Kalai, and David Bau. 2023 · 2023
Cited alongside, same era.
Instructsafety: A unified framework for building multidimensional and explainable safety detector through instruction tuning
Zhexin Zhang, Jiale Cheng, Hao Sun, Jiawen Deng, and Minlie Huang. 2023a · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023 · 2023
Later among the works it cites.
Can we edit factual knowledge by in-context learning?
Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. 2023 · 2023
Later among the works it cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Representation engineering: A top-down approach to AI transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023 · 2023
Later among the works it cites.
Model editing can hurt general abilities of large language models
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng. 2024 · 2024
Closest in time.
Model editing at scale leads to gradual and catastrophic forgetting
Akshat Gupta, Anurag Rao, and Gopala Anumanchipalli. 2024 · 2024
Closest in time.
Sowing the wind, reaping the whirlwind: The impact of editing language models
Rima Hazra, Sayan Layek, Somnath Banerjee, and Soujanya Poria. 2024 · 2024
Closest in time.
Wenyue Hua, Jiang Guo, Mingwen Dong, Henghui Zhu, Patrick Ng, and Zhiguo Wang. 2024 · 2024
Closest in time.
See the unseen: Better context-consistent knowledge-editing by noises
Youcheng Huang, Wenqiang Lei, Zheng Zhang, Jiancheng Lv, and Shuicheng Yan. 2024 · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea. 2024 · 2024
Closest in time.
SWEA: changing factual knowledge in large language models via subject word embedding altering
Xiaopeng Li, Shasha Li, Bin Ji, Shezheng Song, Xi Wang, Jun Ma, Jie Yu, Xiaodong Liu, Jing Wang, and Weimin Zhang. 2024 · 2024
Closest in time.
Large language models relearn removed concepts
Michelle Lo, Shay B. Cohen, and Fazl Barez. 2024 · 2024
Closest in time.
Is it possible to edit large language models robustly?
Xinbei Ma, Tianjie Ju, Jiyang Qiu, Zhuosheng Zhang, Hai Zhao, Lifeng Liu, and Yulong Wang. 2024 · 2024
Closest in time.
MPN: leveraging multilingual patch neuron for cross-lingual model editing
Nianwen Si, Hao Zhang, and Weiqiang Zhang. 2024 · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, and Yue Zhao. 2024 · 2024
Closest in time.
Potential and challenges of model editing for social debiasing
Jianhao Yan, Futing Wang, Yafu Li, and Yue Zhang. 2024 · 2024
Closest in time.
A comprehensive study of knowledge editing for large language models
Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, Siyuan Cheng, Ziwen Xu, Xin Xu, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Lei Liang, Zhiqiang Zhang, Xiaowei Zhu, Jun Zhou, and Huajun Chen. 2024 · 2024
Closest in time.
Prompt-driven LLM safeguarding via directed representation optimization
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. 2024 · 2024
Closest in time.