Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown extraordinary capabilities in understanding and generating text that closely mirrors human communication.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 1911
Earlier work this paper cites.
URL https://api.semanticscholar.org/CorpusID:177285798
Jérôme Seymour Bruner, 1960 · 1960
Earlier work this paper cites.
The course of cognitive growth
Jérôme Seymour Bruner · 1964
Earlier work this paper cites.
Knowledge acquisition - principles and guidelines
Karen L. McGraw and Karan Harbison-Briggs · 1990
Earlier work this paper cites.
What is a knowledge representation?
Randall Davis, Howard E. Shrobe, and Peter Szolovits · 1993
Earlier work this paper cites.
Human-computer interaction: psychology as a science of design
John M Carroll · 1997
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2003
Earlier work this paper cites.
Benjamin Heinzerling and Kentaro Inui · 2008
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2016
Earlier work this paper cites.
Genome-editing technologies for gene and cell therapy
Morgan L Maeder and Charles A Gersbach · 2016
Earlier work this paper cites.
Genome editing comes of age
Jin-Soo Kim · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Expert gate: Lifelong learning with a network of experts
Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Knowledge acquisition–scholarly foundations with knowledge management
N Jayashri and K Kalaiselvi · 2018
Earlier work this paper cites.
Experience replay for continual learning
David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, and Greg Wayne · 2018
Earlier work this paper cites.
Never-ending learning
Tom Mitchell, William Cohen, Estevam Hruschka, Partha Talukdar, Bishan Yang, Justin Betteridge, Andrew Carlson, Bhavana Dalvi, Matt Gardner, Bryan Kisiel, et al · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2018
Earlier work this paper cites.
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Earlier work this paper cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
ERNIE: enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu · 2019
Earlier work this paper cites.
Online continual learning with maximally interfered retrieval
Rahaf Aljundi, Lucas Caccia, Eugene Belilovsky, Massimo Caccia, Min Lin, Laurent Charlin, and Tinne Tuytelaars · 2019
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna M. Wallach, Jennifer T. Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Cem Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
Evaluating commonsense in pre-trained language models
Xuhui Zhou, Yue Zhang, Leyang Cui, and Dandan Huang · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang · 2020
Earlier work this paper cites.
Entities as experts: Sparse memory access with entity supervision
Thibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi, and Tom Kwiatkowski · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Earlier work this paper cites.
Modifying memories in transformer models
Ankit Singh Rawat, Chen Zhu, Daliang Li, Felix Yu, Manzil Zaheer, Sanjiv Kumar, and Srinadh Bhojanapalli · 2020
Earlier work this paper cites.
Editable neural networks
Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitry Pyrkin, Sergei Popov, and Artem Babenko · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Formalizing data deletion in the context of the right to be forgotten
Sanjam Garg, Shafi Goldwasser, and Prashant Nalini Vasudevan · 2020
Earlier work this paper cites.
The promise and challenge of therapeutic genome editing
Jennifer A Doudna · 2020
Earlier work this paper cites.
Knowledgeable machine learning for natural language processing
Xu Han, Zhengyan Zhang, and Zhiyuan Liu · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Earlier work this paper cites.
Can generative pre-trained language models serve as knowledge bases for closed-book qa?
Cunxiang Wang, Pai Liu, and Yue Zhang · 2021
Earlier work this paper cites.
Factual probing is [MASK]: learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen · 2021
Earlier work this paper cites.
Knowledgeable or educated guess? revisiting language models as knowledge bases
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu · 2021
Earlier work this paper cites.
A survey on green deep learning
Jingjing Xu, Wangchunshu Zhou, Zhiyi Fu, Hao Zhou, and Lei Li · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Earlier work this paper cites.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong · 2021
Earlier work this paper cites.
A continual learning survey: Defying forgetting in classification tasks
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Gregory G. Slabaugh, and Tinne Tuytelaars · 2021
Earlier work this paper cites.
Knowledgeable or educated guess? revisiting language models as knowledge bases
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay · 2021
Earlier work this paper cites.
Editing a classifier by rewriting its prediction rules
Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba, and Aleksander Madry · 2021
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq R. Joty, Richard Socher, and Nazneen Fatema Rajani · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Earlier work this paper cites.
Knowledge is power: Symbolic knowledge distillation, commonsense morality, & multimodal script knowledge
Yejin Choi · 2022
Earlier work this paper cites.
ASER: towards large-scale commonsense knowledge acquisition via higher-order selectional preference over eventualities
Hongming Zhang, Xin Liu, Haojie Pan, Haowen Ke, Jiefu Ou, Tianqing Fang, and Yangqiu Song · 2022
Earlier work this paper cites.
Human language understanding & reasoning
Christopher D Manning · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Earlier work this paper cites.
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick · 2022
Earlier work this paper cites.
Can language models serve as temporal knowledge bases?
Ruilin Zhao, Feng Zhao, Guandong Xu, Sixiao Zhang, and Hai Jin · 2022
Earlier work this paper cites.
Time-aware language models as temporal knowledge bases
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen · 2022
Earlier work this paper cites.
A review on language models as knowledge bases
Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona T. Diab, and Marjan Ghazvininejad · 2022
Earlier work this paper cites.
Towards efficient NLP: A standard evaluation and A strong baseline
Xiangyang Liu, Tianxiang Sun, Junliang He, Jiawen Wu, Lingling Wu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, and Xipeng Qiu · 2022
Earlier work this paper cites.
On the impossible safety of large AI models
El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, and John Stephan · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Symbolic knowledge distillation: from general language models to commonsense models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi · 2022
Cited alongside, same era.
A systematic investigation of commonsense knowledge in large language models
Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, and Aida Nematzadeh · 2022
Cited alongside, same era.
A rigorous study of integrated gradients method and extensions to internal neuron attributions
Daniel D Lundstrom, Tianjian Huang, and Meisam Razaviyayn · 2022
Cited alongside, same era.
In-context retrieval-augmented language models
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham · 2023
Later among the works it cites.
Deep class-incremental learning: A survey
Da-Wei Zhou, Qi-Wei Wang, Zhi-Hong Qi, Han-Jia Ye, De-Chuan Zhan, and Ziwei Liu · 2023
Later among the works it cites.
Knowledge unlearning for llms: Tasks, methods, and challenges, 2023
Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang · 2023
Later among the works it cites.
LEACE: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Later among the works it cites.
Unlearn what you want to forget: Efficient unlearning for LLMs
Jiaao Chen and Diyi Yang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finding skill neurons in pre-trained transformer-based language models
Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li · 2022
Cited alongside, same era.
Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction
Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen · 2022
Cited alongside, same era.
PTR: prompt tuning with rules for text classification
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun · 2022
Cited alongside, same era.
Mention memory: incorporating textual knowledge into transformers through entity mention attention
Michiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Fei Sha, and William W. Cohen · 2022
Cited alongside, same era.
Kformer: Knowledge injection in transformer feed-forward layers
Yunzhi Yao, Shaohan Huang, Li Dong, Furu Wei, Huajun Chen, and Ningyu Zhang · 2022
Cited alongside, same era.
Training language models with memory augmentation
Zexuan Zhong, Tao Lei, and Danqi Chen · 2022
Cited alongside, same era.
Decoupling knowledge from memorization: Retrieval-augmented prompt learning
Xiang Chen, Lei Li, Ningyu Zhang, Xiaozhuan Liang, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen · 2022
Cited alongside, same era.
Polyglot or not? measuring multilingual encyclopedic knowledge in foundation models
Tim Schott, Daniel Furman, and Shreshta Bhat · 2023
Later among the works it cites.
BertNet: Harvesting knowledge graphs with arbitrary relations from pretrained language models
Shibo Hao, Bowen Tan, Kaiwen Tang, Bin Ni, Xiyan Shao, Hengzhe Zhang, Eric Xing, and Zhiting Hu · 2023
Later among the works it cites.
Editing commonsense knowledge in gpt, 2023
Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao, Xiang Lorraine Li, Sarah Wiegreffe, and Niket Tandon · 2023
Later among the works it cites.
MQuAKE: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher Manning, Christopher Potts, and Danqi Chen · 2023
Later among the works it cites.
Can we edit factual knowledge by in-context learning?
Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang · 2023
Later among the works it cites.
Evaluating the ripple effects of knowledge editing in language models, 2023
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva · 2023
Later among the works it cites.
Methods for measuring, updating, and visualizing factual beliefs in language models
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer · 2023
Later among the works it cites.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau · 2023
Later among the works it cites.
Massive editing for large language models via meta learning
Chenmien Tan, Ge Zhang, and Jie Fu · 2023
Later among the works it cites.
Untying the reversal curse via bidirectional language model editing, 2023
Jun-Yu Ma, Jia-Chen Gu, Zhen-Hua Ling, Quan Liu, and Cong Liu · 2023
Later among the works it cites.
Self-knowledge guided retrieval augmentation for large language models
Yile Wang, Peng Li, Maosong Sun, and Yang Liu · 2023
Later among the works it cites.
Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts, 2023
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Qinan Yu, Jack Merullo, and Ellie Pavlick · 2023
Later among the works it cites.
Shiwen Ni, Dingwei Chen, Chengming Li, Xiping Hu, Ruifeng Xu, and Min Yang · 2023
Later among the works it cites.
A divide and conquer framework for knowledge editing
Xiaoqi Han, Ru Li, Xiaoli Li, and Jeff Z. Pan · 2023
Later among the works it cites.
The reversal curse: Llms trained on ”a is b” fail to learn ”b is a”, 2023
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2023
Later among the works it cites.
Emptying the ocean with a spoon: Should we edit models?
Yuval Pinter and Michael Elhadad · 2023
Later among the works it cites.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales · 2023
Later among the works it cites.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Knowledge sanitization of large language models
Yoichi Ishibashi and Hidetoshi Shimodaira · 2023
Later among the works it cites.
Evaluating dependencies in fact editing for language models: Specificity and implication awareness
Zichao Li, Ines Arous, Siva Reddy, and Jackie Cheung · 2023
Later among the works it cites.
Language anisotropic cross-lingual model editing
Yang Xu, Yutai Hou, Wanxiang Che, and Min Zhang · 2023
Later among the works it cites.
Can LMs learn new entities from descriptions? challenges in propagating injected knowledge
Yasumasa Onoe, Michael Zhang, Shankar Padmanabhan, Greg Durrett, and Eunsol Choi · 2023
Later among the works it cites.
Assessing knowledge editing in language models via relation perspective
Yifan Wei, Xiaoyan Yu, Huanhuan Ma, Fangyu Lei, Yixuan Weng, Ran Song, and Kang Liu · 2023
Later among the works it cites.
Dune: Dataset for unified editing
Afra Feyza Akyürek, Eric Pan, Garry Kuwanto, and Derry Wijaya · 2023
Later among the works it cites.
Opencompass: A universal evaluation platform for foundation models
OpenCompass Contributors · 2023
Later among the works it cites.
Improving sequential model editing with fact retrieval
Xiaoqi Han, Ru Li, Hongye Tan, Wang Yuanlong, Qinghua Chai, and Jeff Pan · 2023
Later among the works it cites.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2023
Later among the works it cites.
Do localization methods actually localize memorized data in llms?, 2023
Ting-Yun Chang, Jesse Thomason, and Robin Jia · 2023
Later among the works it cites.
Klob: a benchmark for assessing knowledge locating methods in language models, 2023
Yiming Ju and Zheng Zhang · 2023
Later among the works it cites.
Propagating knowledge updates to lms through distillation, 2023
Shankar Padmanabhan, Yasumasa Onoe, Michael J. Q. Zhang, Greg Durrett, and Eunsol Choi · 2023
Later among the works it cites.
Dola: Decoding by contrasting layers improves factuality in large language models, 2023
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He · 2023
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Later among the works it cites.
Edit at your own risk: evaluating the robustness of edited models to distribution shifts, 2023
Davis Brown, Charles Godfrey, Cody Nizinski, Jonathan Tu, and Henry Kvinge · 2023
Later among the works it cites.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard · 2023
Later among the works it cites.
Discovering knowledge-critical subnetworks in pretrained language models
Deniz Bayazit, Negar Foroutan, Zeming Chen, Gail Weiss, and Antoine Bosselut · 2023
Later among the works it cites.
An empirical study of multimodal model merging
Yi-Lin Sung, Linjie Li, Kevin Lin, Zhe Gan, Mohit Bansal, and Lijuan Wang · 2023
Later among the works it cites.
Ablating concepts in text-to-image diffusion models
Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu · 2023
Later among the works it cites.
Localizing and editing knowledge in text-to-image generative models, 2023
Samyadeep Basu, Nanxuan Zhao, Vlad Morariu, Soheil Feizi, and Varun Manjunatha · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang and Maneesh Agrawala · 2023
Later among the works it cites.
Refact: Updating text-to-image models by editing the text encoder, 2023
Dana Arad, Hadas Orgad, and Yonatan Belinkov · 2023
Later among the works it cites.
Erasing concepts from diffusion models, 2023
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau · 2023
Later among the works it cites.
AI alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao · 2023
Later among the works it cites.
Provenance documentation to enable explainable and trustworthy AI: A literature review
Amruta Kale, Tin Nguyen, Frederick C. Harris Jr., Chenhao Li, Jiyin Zhang, and Xiaogang Ma · 2023
Later among the works it cites.
Unveiling the implicit toxicity in large language models
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang · 2023
Later among the works it cites.
arXiv preprint arXiv:2312.14302 , 2023
Exploiting novel gpt-4 apis · 2023
Later among the works it cites.
Recent advances towards safe, responsible, and moral dialogue systems: A survey
Jiawen Deng, Hao Sun, Zhexin Zhang, Jiale Cheng, and Minlie Huang · 2023
Later among the works it cites.
Xinshuo Hu, Dongfang Li, Zihao Zheng, Zhenyu Liu, Baotian Hu, and Min Zhang · 2023
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed Chi, Denny Zhou, and Jason Wei · 2023
Later among the works it cites.
Layered bias: Interpreting bias in pretrained large language models
Nirmalendu Prakash and Roy Ka-Wei Lee · 2023
Later among the works it cites.
Unlearning bias in language models by partitioning gradients
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji · 2023
Later among the works it cites.
Debiasing algorithm through model adaptation, 2023
Tomasz Limisiewicz, David Mareček, and Tomáš Musil · 2023
Later among the works it cites.
Privacy issues in large language models: A survey
Seth Neel and Peter Chang · 2023
Later among the works it cites.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Later among the works it cites.
Lamp: When large language models meet personalization
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani · 2023
Later among the works it cites.
Do llms possess a personality? making the MBTI test an amazing evaluation for large language models
Keyu Pan and Yawen Zeng · 2023
Later among the works it cites.
Personality traits in large language models
Mustafa Safdari, Greg Serapio-García, Clément Crepy, Stephen Fitz, Peter Romero, Luning Sun, Marwa Abdulhai, Aleksandra Faust, and Maja Mataric · 2023
Later among the works it cites.
Characterchat: Learning towards conversational AI with personalized social support
Quan Tu, Chuanqi Chen, Jinpeng Li, Yanran Li, Shuo Shang, Dongyan Zhao, Ran Wang, and Rui Yan · 2023
Later among the works it cites.
Shengyu Mao, Ningyu Zhang, Xiaohan Wang, Mengru Wang, Yunzhi Yao, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen · 2023
Later among the works it cites.
Emotionally numb or empathetic? evaluating how llms feel using emotionbench, 2023
Jen tse Huang, Man Ho Lam, Eric John Li, Shujie Ren, Wenxuan Wang, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu · 2023
Later among the works it cites.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Prasad Majumder, and Niket Tandon · 2023
Later among the works it cites.
Pmet: Precise model editing in a transformer
Xiaopeng Li, Shasha Li, Shezheng Song, Jing Yang, Jun Ma, and Jie Yu · 2024
Closest in time.
Alphaedit: Null-space constrained knowledge editing for language models
Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Xiang Wang, Xiangnan He, and Tat-seng Chua · 2024
Closest in time.
Editing language model-based knowledge graph embeddings
Siyuan Cheng, Ningyu Zhang, Bozhong Tian, Zelin Dai, Feiyu Xiong, Wei Guo, and Huajun Chen · 2024
Closest in time.