Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) suffer from hallucinations, referring to the non-factual information in generated content, despite their superior capacities across tasks.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar · 2012
Earlier work this paper cites.
Fake news detection on social media: A data mining perspective
Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu · 2017
Earlier work this paper cites.
Combining interventions to reduce the spread of viral misinformation
Joseph B Bak-Coleman, Ian Kennedy, Morgan Wack, Andrew Beers, Joseph S Schafer, Emma S Spiro, Kate Starbird, and Jevin D West · 2022
Earlier work this paper cites.
Canyu Chen, Haoran Wang, Matthew Shapiro, Yunyu Xiao, Fei Wang, and Kai Shu · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Memory-based model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, and Chelsea Finn · 2022
Earlier work this paper cites.
Reviewing interventions to address misinformation: the need to expand our vision beyond an individualistic focus
Zhila Aghajari, Eric PS Baumer, and Dominic DiFranzo · 2023
Earlier work this paper cites.
Dune: Dataset for unified editing
Afra Feyza Akyürek, Eric Pan, Garry Kuwanto, and Derry Wijaya · 2023
Earlier work this paper cites.
Pokemqa: Programmable knowledge editing for multi-hop question answering
Hengrui Gu, Kaixiong Zhou, Xiaotian Han, Ninghao Liu, Ruobing Wang, and Xin Wang · 2023
Earlier work this paper cites.
Reinforcement learning-based counter-misinformation response generation: a case study of covid-19 vaccine misinformation
Bing He, Mustaque Ahamad, and Srijan Kumar · 2023
Earlier work this paper cites.
Detecting edit failures in large language models: An improved specificity benchmark
Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, and Fazl Barez · 2023
Earlier work this paper cites.
Evaluating dependencies in fact editing for language models: Specificity and implication awareness
Zichao Li, Ines Arous, Siva Reddy, and Jackie Chi Kit Cheung · 2023
Earlier work this paper cites.
Untying the reversal curse via bidirectional language model editing
Jun-Yu Ma, Jia-Chen Gu, Zhen-Hua Ling, Quan Liu, and Cong Liu · 2023
Earlier work this paper cites.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau · 2023
Earlier work this paper cites.
Exploiting user comments for early detection of fake news prior to users’ commenting
Qiong Nan, Qiang Sheng, Juan Cao, Yongchun Zhu, Danding Wang, Guang Yang, Jintao Li, and Kai Shu · 2023
Earlier work this paper cites.
Evaluating the social impact of generative ai systems in systems and society
Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Canyu Chen, Hal Daumé III, Jesse Dodge, Isabella Duan, et al · 2023
Earlier work this paper cites.
Attacking fake news detectors via manipulating news social engagement
Haoran Wang, Yingtong Dou, Canyu Chen, Lichao Sun, Philip S Yu, and Kai Shu · 2023
Earlier work this paper cites.
Assessing knowledge editing in language models via relation perspective
Yifan Wei, Xiaoyan Yu, Huanhuan Ma, Fangyu Lei, Yixuan Weng, Ran Song, and Kang Liu · 2023
Earlier work this paper cites.
Eva-kellm: A new benchmark for evaluating knowledge editing of llms
Suhang Wu, Minlong Peng, Yue Chen, Jinsong Su, and Mingming Sun · 2023
Earlier work this paper cites.
Editing large language models: Problems, methods, and opportunities
Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang · 2023
Earlier work this paper cites.
Melo: Enhancing model editing with neuron-indexed dynamic lora
Lang Yu, Qin Chen, Jie Zhou, and Liang He · 2023
Earlier work this paper cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi · 2023
Earlier work this paper cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen · 2023
Earlier work this paper cites.
Can we edit factual knowledge by in-context learning?
Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang · 2023
Earlier work this paper cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen · 2023
Earlier work this paper cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Earlier work this paper cites.
Alimohammad Beigi, Zhen Tan, Nivedh Mudiam, Canyu Chen, Kai Shu, and Huan Liu · 2024
Cited alongside, same era.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva · 2024
Cited alongside, same era.
Unke: Unstructured knowledge editing in large language models
Jingcheng Deng, Zihao Wei, Liang Pang, Hanxing Ding, Huawei Shen, and Xueqi Cheng · 2024
Cited alongside, same era.
Alphaedit: Null-space constrained knowledge editing for language models
Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Xiang Wang, Xiangnan He, and Tat-seng Chua · 2024
Cited alongside, same era.
Retrieval meets reasoning: Dynamic in-context editing for long-text understanding
Detox: Toxic subspace projection for model editing
Rheeya Uppaal, Apratim De, Yiting He, Yiquao Zhong, and Junjie Hu · 2024
Closest in time.
Introducing v0. 5 of the ai safety benchmark from mlcommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M Ahmed, Victor Akinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, et al · 2024
Closest in time.
Updating language models with unstructured facts: Towards practical knowledge editing
Xiaobao Wu, Liangming Pan, William Yang Wang, and Anh Tuan Luu · 2024
Closest in time.
Memla: Enhancing multilingual knowledge editing with neuron-masked low-rank adaptation
Jiakuan Xie, Pengfei Cao, Yuheng Chen, Yubo Chen, Kang Liu, and Jun Zhao · 2024
Closest in time.
Editing factual knowledge and explanatory ability of medical large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weizhi Fei, Xueyan Niu, Guoqing Xie, Yanhua Zhang, Bo Bai, Lei Deng, and Wei Han · 2024
Cited alongside, same era.
A primer on the inner workings of transformer-based language models
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R Costa-jussà · 2024
Cited alongside, same era.
Model editing by pure fine-tuning
Govind Gangadhar and Karl Stratos · 2024
Cited alongside, same era.
Concept-rot: Poisoning concepts in large language models with model editing
Keltin Grimes, Marco Christiani, David Shriver, and Marissa Connor · 2024
Cited alongside, same era.
Model editing harms general abilities of large language models: Regularization to the rescue
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng · 2024
Cited alongside, same era.
Model editing at scale leads to gradual and catastrophic forgetting
Akshat Gupta, Anurag Rao, and Gopala Anumanchipalli · 2024
Cited alongside, same era.
Aging with grace: Lifelong model editing with discrete key-value adaptors
Tom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim, and Marzyeh Ghassemi · 2024
Cited alongside, same era.
The landscape of user-centered misinformation interventions-a systematic literature review
Katrin Hartwig, Frederic Doell, and Christian Reuter · 2024
Cited alongside, same era.
Derong Xu, Ziheng Zhang, Zhihong Zhu, Zhenxi Lin, Qidong Liu, Xian Wu, Tong Xu, Xiangyu Zhao, Yefeng Zheng, and Enhong Chen · 2024
Closest in time.
Potential and challenges of model editing for social debiasing
Jianhao Yan, Futing Wang, Yafu Li, and Yue Zhang · 2024
Closest in time.
The butterfly effect of model editing: Few edits can trigger large language models collapse
Wanli Yang, Fei Sun, Xinyu Ma, Xun Liu, Dawei Yin, and Xueqi Cheng · 2024
Closest in time.
The fall of rome: Understanding the collapse of llms in model editing
Wanli Yang, Fei Sun, Jiajun Tan, Xinyu Ma, Du Su, Dawei Yin, and Huawei Shen · 2024
Closest in time.
Knowledge circuits in pretrained transformers
Yunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang, Ziwen Xu, Shumin Deng, and Huajun Chen · 2024
Closest in time.
History matters: Temporal knowledge editing in large language model
Xunjian Yin, Jin Jiang, Liming Yang, and Xiaojun Wan · 2024
Closest in time.
Has this fact been edited? detecting knowledge edits in language models
Paul Youssef, Zhixue Zhao, Christin Seifert, and Jörg Schlötterer · 2024
Closest in time.
Evidence-driven retrieval augmented response generation for online misinformation
Zhenrui Yue, Huimin Zeng, Yimeng Lu, Lanyu Shang, Yang Zhang, and Dong Wang · 2024
Closest in time.
Visual-oriented fine-grained knowledge editing for multimodal large language models
Zhen Zeng, Leijiang Gu, Xun Yang, Zhangling Duan, Zenglin Shi, and Meng Wang · 2024
Closest in time.
Trustworthiness in retrieval-augmented generation systems: A survey
Yujia Zhou, Yan Liu, Xiaoxi Li, Jiajie Jin, Hongjin Qian, Zheng Liu, Chaozhuo Li, Zhicheng Dou, Tsung-Yi Ho, and Philip S Yu · 2024
Closest in time.
MMKE-bench: A multimodal editing benchmark for diverse visual knowledge
Yuntao Du., Kailin Jiang, Zhi Gao, Chenrui Shi, Zilong Zheng, Siyuan Qi, and Qing Li · 2025
Closest in time.
Editing large language models via adaptive gradient guidance
Xiaojie Gu, Guangxu Chen, Shuliang Liu, Jungang Li, Aiwei Liu, Sicheng Tao, Junyan Zhang, and Xuming Hu · 2025
Closest in time.
Anyedit: Edit any knowledge encoded in language models
Houcheng Jiang, Junfeng Fang, Ningyu Zhang, Guojun Ma, Mingyang Wan, Xiang Wang, Xiangnan He, and Tat-seng Chua · 2025
Closest in time.
Reinforced lifelong editing for language models
Zherui Li, Houcheng Jiang, Hao Chen, Baolong Bi, Zhenhong Zhou, Fei Sun, Junfeng Fang, and Xiang Wang · 2025
Closest in time.
Mitigating heterogeneous token overfitting in llm knowledge editing
Tianci Liu, Zihan Dong, Linjun Zhang, Haoyu Wang, and Jing Gao · 2025
Closest in time.
Towards trustworthy retrieval augmented generation for large language models: A survey
Bo Ni, Zheyuan Liu, Leyao Wang, Yongjia Lei, Yuying Zhao, Xueqi Cheng, Qingkai Zeng, Luna Dong, Yinglong Xia, Krishnaram Kenthapadi, et al · 2025
Closest in time.
Searchrag: Can search engines be helpful for llm-based medical question answering?
Yucheng Shi, Tianze Yang, Canyu Chen, Quanzheng Li, Tianming Liu, Xiang Li, and Ninghao Liu · 2025
Closest in time.
Edit once, update everywhere: A simple framework for cross-lingual knowledge synchronization in llms
Yuchen Wu, Liang Ding, Li Shen, and Dacheng Tao · 2025
Closest in time.
The mirage of model editing: Revisiting evaluation in the wild
Wanli Yang, Fei Sun, Jiajun Tan, Xinyu Ma, Qi Cao, Dawei Yin, Huawei Shen, and Xueqi Cheng · 2025
Closest in time.
Position: Editing large language models poses serious safety risks
Paul Youssef, Zhixue Zhao, Daniel Braun, Jörg Schlötterer, and Christin Seifert · 2025
Closest in time.
Explainable and efficient editing for large language models
Tianyu Zhang, Junfeng Fang, Houcheng Jiang, Baolong Bi, Xiang Wang, and Xiangnan He · 2025
Closest in time.
Fleke: Federated locate-then-edit knowledge editing
Zongkai Zhao, Guozeng Xu, Xiuhua Li, Kaiwen Wei, and Jiang Zhong · 2025
Closest in time.