Fetching the paper…
Reading the bibliography…
Model editing aims to efficiently alter the behavior of Large Language Models (LLMs) within a desired scope, while ensuring no adverse impact on other inputs.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 1910
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
B. Gliwa, I. Mochol, M. Biesek, and A. Wawer · 1911
Earlier work this paper cites.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
E. F. T. K. Sang and F. D. Meulder · 2003
Earlier work this paper cites.
Machine Learning Challenges, Evaluating Predictive Uncertainty, Visual Object Classification and Recognizing Textual Entailment, First PASCAL Machine Learning Challenges Workshop, MLCW 2005, Southampton, UK, April 11-13, 2005, Revised Selected Papers , volume 3944 of Lecture Notes in Computer Science , 2006. Springer
J. Q. Candela, I. Dagan, B. Magnini, and F. d’Alché-Buc, editors · 2006
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
D. Eigen, M. Ranzato, and I. Sutskever · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
The ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems
R. Lowe, N. Pow, I. Serban, and J. Pineau · 2015
Earlier work this paper cites.
Reading wikipedia to answer open-domain questions
D. Chen, A. Fisch, J. Weston, and A. Bordes · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
O. Levy, M. Seo, E. Choi, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. V. Le, G. E. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. P. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
K. Lee, M. Chang, and K. Toutanova · 2019
Earlier work this paper cites.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. S. H. Lewis, A. Bakhtin, Y. Wu, and A. H. Miller · 2019
Earlier work this paper cites.
Mutual: A dataset for multi-turn dialogue reasoning
L. Cui, Y. Wu, S. Liu, Y. Zhang, and M. Zhou · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
H. Hazimeh, Z. Zhao, A. Chowdhery, M. Sathiamoorthy, Y. Chen, R. Mazumder, L. Hong, and E. H. Chi · 2021
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2021
Cited alongside, same era.
BASE layers: Simplifying training of large, sparse models
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. S. Pinto, D. Keysers, and N. Houlsby · 2021
Cited alongside, same era.
Towards understanding the mixture-of-experts layer in deep learning
Z. Chen, Y. Deng, Y. Wu, Q. Gu, and Y. Li · 2022
Cited alongside, same era.
Knowledge neurons in pretrained transformers
D. Dai, L. Dong, Y. Hao, Z. Sui, B. Chang, and F. Wei · 2022
Cited alongside, same era.
Calibrating factual knowledge in pretrained language models
Q. Dong, D. Dai, Y. Song, J. Xu, Z. Sui, and L. Li · 2022
Transformer-patcher: One mistake worth one neuron
Z. Huang, Y. Shen, X. Zhang, J. Zhou, W. Rong, and Z. Xiong · 2023
Later among the works it cites.
Large language models with controllable working memory
D. Li, A. S. Rawat, M. Zaheer, X. Wang, M. Lukasik, A. Veit, F. X. Yu, and S. Kumar · 2023
Later among the works it cites.
Mass-editing memory in a transformer
K. Meng, A. S. Sharma, A. J. Andonian, Y. Belinkov, and D. Bau · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Emptying the ocean with a spoon: Should we edit models?
Y. Pinter and M. Elhadad · 2023
Later among the works it cites.
Mixture-of-experts meets instruction tuning: A winning combination for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. P. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. S. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Memory-assisted prompt editing to improve GPT-3 after deployment
A. Madaan, N. Tandon, P. Clark, and Y. Yang · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Cited alongside, same era.
Fast model editing at scale
E. Mitchell, C. Lin, A. Bosselut, C. Finn, and C. D. Manning · 2022
Cited alongside, same era.
S. Shen, L. Hou, Y. Zhou, N. Du, S. Longpre, J. Wei, H. W. Chung, B. Zoph, W. Fedus, X. Chen, et al · 2023
Later among the works it cites.
Moec: Mixture of expert clusters
Y. Xie, S. Huang, T. Chen, and F. Wei · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Y. Yao, P. Wang, B. Tian, S. Cheng, Z. Li, S. Deng, H. Chen, and N. Zhang · 2023
Later among the works it cites.
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning
T. Zadouri, A. Üstün, A. Ahmadian, B. Ermis, A. Locatelli, and S. Hooker · 2023
Later among the works it cites.
Can we edit factual knowledge by in-context learning?
C. Zheng, L. Li, Q. Dong, Y. Fan, Z. Wu, J. Xu, and B. Chang · 2023
Later among the works it cites.
Higher layers need more lora experts
C. Gao, K. Chen, J. Rao, B. Sun, R. Liu, D. Peng, Y. Zhang, X. Guo, J. Yang, and V. S. Subrahmanian · 2024
Closest in time.
Model editing can hurt general abilities of large language models
J. Gu, H. Xu, J. Ma, P. Lu, Z. Ling, K. Chang, and N. Peng · 2024
Closest in time.
Model editing at scale leads to gradual and catastrophic forgetting
A. Gupta, A. Rao, and G. Anumanchipalli · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de Las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
Consecutive model editing with batch alongside hook layers
S. Li, Y. Deng, D. Cai, H. Lu, L. Chen, and W. Lam · 2024
Closest in time.
R. Wang and P. Li · 2024
Closest in time.
Scalable model editing via customized expert networks
Z. Yao, Y. He, T. Qi, and M. Li · 2024
Closest in time.
A comprehensive study of knowledge editing for large language models
N. Zhang, Y. Yao, B. Tian, P. Wang, S. Deng, M. Wang, Z. Xi, S. Mao, J. Zhang, Y. Ni, S. Cheng, Z. Xu, X. Xu, J. Gu, Y. Jiang, P. Xie, F. Huang, L. Liang, Z. Zhang, X. Zhu, J. Zhou, and H. Chen · 2024
Closest in time.