Fetching the paper…
Reading the bibliography…
Knowledge editing methods (KEs) can update language models' obsolete or inaccurate knowledge learned from pre-training.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire. 1997 · 1997
Earlier work this paper cites.
Pattern classification, chapter nonparametric techniques
RO Duda, PE Hart, DG Stork, and Alexandru Ionescu. 2000 · 2000
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Backdoor attacks on pre-trained models by layerwise weight poisoning
Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruotian Ma, and Xipeng Qiu. 2021 · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Earlier work this paper cites.
A comparative study of using pre-trained language models for toxic comment classification
Zhixue Zhao, Ziqi Zhang, and Frank Hopfgartner. 2021 · 2021
Earlier work this paper cites.
Transformer-patcher: One mistake worth one neuron
Zeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou, Wenge Rong, and Zhang Xiong. 2022 · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Stealthy and persistent unalignment on large language models via backdoor injections
Yuanpu Cao, Bochuan Cao, and Jinghui Chen. 2023 · 2023
Earlier work this paper cites.
Decoding the threat landscape: ChatGPT, FraudGPT, and WormGPT in social engineering attacks
Polra Victor Falade. 2023 · 2023
Earlier work this paper cites.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023 · 2023
Earlier work this paper cites.
Controllable text generation via probability density estimation in the latent space
Yuxuan Gu, Xiaocheng Feng, Sicheng Ma, Lingyuan Zhang, Heng Gong, Weihong Zhong, and Bing Qin. 2023 · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark. 2023 · 2023
Cited alongside, same era.
Aging with grace: Lifelong model editing with discrete key-value adaptors
Thomas Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim, and Marzyeh Ghassemi. 2023 · 2023
Cited alongside, same era.
Inspecting and editing knowledge representations in language models
Evan Hernandez, Belinda Z Li, and Jacob Andreas. 2023 · 2023
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. 2023 · 2023
Cited alongside, same era.
Unlearning bias in language models by partitioning gradients
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. 2023a · 2023
Later among the works it cites.
Can we edit factual knowledge by in-context learning?
Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. 2023 · 2023
Later among the works it cites.
Can editing LLMs inject harm?
Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen, Shiyang Lai, Xiongxiao Xu, Jia-Chen Gu, Jindong Gu, Huaxiu Yao, Chaowei Xiao, Xifeng Yan, William Yang Wang, Philip Torr, Dawn Song, and Kai Shu. 2024 · 2024
Closest in time.
Universal neurons in GPT2 language models
Wes Gurnee, Theo Horsley, Zifan Carl Guo, Tara Rezaei Kheirkhah, Qinyi Sun, Will Hathaway, Neel Nanda, and Dimitris Bertsimas. 2024 · 2024
Closest in time.
" flex tape can’t fix that": Bias and misinformation in edited language models
Karina Halevy, Anna Sotnikova, Badr AlKhamissi, Syrielle Montariol, and Antoine Bosselut. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models
Aly Kassem, Omar Mahmoud, and Sherif Saad. 2023 · 2023
Cited alongside, same era.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Cited alongside, same era.
Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A Rossi, and Thien Huu Nguyen. 2023 · 2023
Cited alongside, same era.
A survey on knowledge editing of neural networks
Vittorio Mazzia, Alessandro Pedrani, Andrea Caciolai, Kay Rottmann, and Davide Bernardi. 2023 · 2023
Cited alongside, same era.
Mass editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2023 · 2023
Cited alongside, same era.
On the risk of misinformation pollution with large language models
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Wang. 2023 · 2023
Cited alongside, same era.
Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt
Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun. 2023 · 2023
Cited alongside, same era.
Closest in time.
Flooding spread of manipulated knowledge in llm-based multi-agent communities
Tianjie Ju, Yiting Wang, Xinbei Ma, Pengzhou Cheng, Haodong Zhao, Yulong Wang, Lifeng Liu, Jian Xie, Zhuosheng Zhang, and Gongshen Liu. 2024 · 2024
Closest in time.
Lingbo Mo, Boshi Wang, Muhao Chen, and Huan Sun. 2024 · 2024
Closest in time.
What does the knowledge neuron thesis have to do with knowledge?
Jingcheng Niu, Andrew Liu, Zining Zhu, and Gerald Penn. 2024 · 2024
Closest in time.
Can sensitive information be deleted from LLMs? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal. 2024 · 2024
Closest in time.
Megen: Generative backdoor in large language models via model editing
Jiyang Qiu, Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2024 · 2024
Closest in time.
Long-form evaluation of model editing
Domenic Rosati, Robie Gonzales, Jinkun Chen, Xuemin Yu, Yahya Kayani, Frank Rudzicz, and Hassan Sajjad. 2024 · 2024
Closest in time.
Massive editing for large language models via meta learning
Chenmien Tan, Ge Zhang, and Jie Fu. 2024 · 2024
Closest in time.
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
Peng Wang, Ningyu Zhang, Bozhong Tian, Zekun Xi, Yunzhi Yao, Ziwen Xu, Mengru Wang, Shengyu Mao, Xiaohan Wang, Siyuan Cheng, Kangwei Liu, Yuansheng Ni, Guozhou Zheng, and Huajun Chen. 2024 · 2024
Closest in time.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024 · 2024
Closest in time.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024 · 2024
Closest in time.
The Queen of England is not England‘s Queen: On the Lack of Factual Coherency in PLMs
Paul Youssef, Jörg Schlötterer, and Christin Seifert. 2024 · 2024
Closest in time.
Position: Editing Large Language Models Poses Serious Safety Risks
Paul Youssef, Zhixue Zhao, Daniel Braun, Jörg Schlötterer, and Christin Seifert. 2025 · 2025
Closest in time.