Fetching the paper…
Reading the bibliography…
The model editing problem concerns how language models should learn new facts about the world over time.
Studies in the logic of confirmation (i.)
Carl G Hempel. 1945 · 1945
Earlier work this paper cites.
Rational choice and the structure of the environment
Herbert A Simon. 1956 · 1956
Earlier work this paper cites.
The web of belief
Willard V Quine and Joseph Silbert Ullian. 1970 · 1970
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty
Amos Tversky and Daniel Kahneman. 1974 · 1974
Earlier work this paper cites.
Counterfactual dependence and time’s arrow
David Lewis. 1979 · 1979
Earlier work this paper cites.
On the logic of theory change: Partial meet contraction and revision functions
Carlos E. Alchourrón, Peter Gärdenfors, and David Makinson. 1985 · 1985
Earlier work this paper cites.
How Bayesian confirmation theory handles the paradox of the ravens
Branden Fitelson and James Hawthorne. 2010 · 2010
Earlier work this paper cites.
What can we learn from collective human opinions on natural language inference data?
Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020 · 2010
Earlier work this paper cites.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar. 2020 · 2012
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky. 2015 · 2015
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2021 · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Earlier work this paper cites.
Speech Acts
Mitchell Green. 2021 · 2021
Earlier work this paper cites.
Do language models have beliefs? methods for detecting, updating, and visualizing model beliefs
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. 2021 · 2021
Earlier work this paper cites.
Mind the gap: Assessing temporal generalization in neural language models
Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomas Kocisky, Sebastian Ruder, et al. 2021 · 2021
Earlier work this paper cites.
Entity-based knowledge conflicts in question answering
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021 · 2021
Earlier work this paper cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2021 · 2021
Earlier work this paper cites.
Learning from disagreement: A survey
Alexandra N Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021 · 2021
Earlier work this paper cites.
Kepler: A unified model for knowledge embedding and pre-trained language representation
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021 · 2021
Earlier work this paper cites.
Distributed nli: Learning to predict human opinion distributions for language reasoning
Xiang Zhou, Yixin Nie, and Mohit Bansal. 2021 · 2021
Earlier work this paper cites.
A review on language models as knowledge bases
Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad. 2022 · 2022
Earlier work this paper cites.
Language models as agent models
Jacob Andreas. 2022 · 2022
Cited alongside, same era.
Time-Aware Language Models as Temporal Knowledge Bases
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022 · 2022
Cited alongside, same era.
Logic of Belief Revision
Sven Ove Hansson. 2022 · 2022
Cited alongside, same era.
Peft: State-of-the-art parameter-efficient fine-tuning methods
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. 2022 · 2022
Cited alongside, same era.
Locating and editing factual knowledge in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Can lms learn new entities from descriptions? challenges in propagating injected knowledge
Yasumasa Onoe, Michael JQ Zhang, Shankar Padmanabhan, Greg Durrett, and Eunsol Choi. 2023 · 2023
Later among the works it cites.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
Emptying the ocean with a spoon: Should we edit models?
Yuval Pinter and Michael Elhadad. 2023 · 2023
Later among the works it cites.
Why think step by step? reasoning emerges from the locality of experience
Ben Prystawski, Michael Li, and Noah Goodman. 2023 · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
The’problem’of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Cited alongside, same era.
Counterfactuals
W. Starr. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Probabilistic coherence, logical consistency, and bayesian learning: Neural language models as epistemic agents
Gregor Betz and Kyle Richardson. 2023 · 2023
Cited alongside, same era.
Edit at your own risk: evaluating the robustness of edited models to distribution shifts
Davis Brown, Charles Godfrey, Cody Nizinski, Jonathan Tu, and Henry Kvinge. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2023 · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al. 2023 · 2023
Later among the works it cites.
Assessing knowledge editing in language models via relation perspective
Yifan Wei, Xiaoyan Yu, Huanhuan Ma, Fangyu Lei, Yixuan Weng, Ran Song, and Kang Liu. 2023 · 2023
Later among the works it cites.
Depn: Detecting and editing privacy neurons in pretrained language models
Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. 2023 · 2023
Later among the works it cites.
Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2023 · 2023
Later among the works it cites.
Rongwu Xu, Brian S Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu. 2023 · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang. 2023 · 2023
Later among the works it cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2024 · 2024
Closest in time.
Context versus prior knowledge in language models
Kevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer C White, Aaron Schein, and Ryan Cotterell. 2024 · 2024
Closest in time.
Mulan: A study of fact mutability in language models
Constanza Fierro, Nicolas Garneau, Emanuele Bugliarello, Yova Kementchedjhieva, and Anders Søgaard. 2024 · 2024
Closest in time.
Model editing by pure fine-tuning
Govind Gangadhar and Karl Stratos. 2024 · 2024
Closest in time.
Are language models rational? the case of coherence norms and belief revision
Thomas Hofweber, Peter Hase, Elias Stengel-Eskin, and Mohit Bansal. 2024 · 2024
Closest in time.
Wenyue Hua, Jiang Guo, Mingwen Dong, Henghui Zhu, Patrick Ng, and Zhiguo Wang. 2024 · 2024
Closest in time.
Personas as a way to model truthfulness in language models
Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, and He He. 2024 · 2024
Closest in time.
Implicit meta-learning may lead language models to trust more reliable sources
Dmitrii Krasheninnikov, Egor Krasheninnikov, Bruno Mlodozeniec, Tegan Maharaj, and David Krueger. 2024 · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al. 2024 · 2024
Closest in time.
Bayesian Epistemology
Hanti Lin. 2024 · 2024
Closest in time.
Taxi: Evaluating categorical knowledge editing for language models
Derek Powell, Walter Gerych, and Thomas Hartvigsen. 2024 · 2024
Closest in time.
What evidence do language models find convincing?
Alexander Wan, Eric Wallace, and Dan Klein. 2024 · 2024
Closest in time.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty
Kaitlyn Zhou, Jena D Hwang, Xiang Ren, and Maarten Sap. 2024 · 2024
Closest in time.