Fetching the paper…
Reading the bibliography…
Large language models learn and continually learn through the accumulation of gradient-based updates, but how individual pieces of new information affect existing knowledge, leading to both beneficial generalization and problematic hallucination, remains poorly understood.
Differentially Private Learning with Adaptive Clipping
G. Andrew, O. Thakkar, H. B. McMahan, and S. Ramaswamy · 1905
Earlier work this paper cites.
What Do Compressed Deep Neural Networks Forget?
S. Hooker, A. Courville, G. Clark, Y. Dauphin, and A. Frome · 1911
Earlier work this paper cites.
Facilitation in recognizing pairs of words: Evidence of a dependence between retrieval operations
D. Meyer and R. Schvaneveldt · 1971
Earlier work this paper cites.
Priming effects in word-fragment completion are independent of recognition memory
E. Tulving, D. Schacter, and H. Stark · 1982
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory
J. K. McClelland, B. K. McNaughton, and R. O’Reilly · 1995
Earlier work this paper cites.
How Much Knowledge Can You Pack Into the Parameters of a Language Model?
A. Roberts, C. Raffel, and N. Shazeer · 2002
Earlier work this paper cites.
Weight Poisoning Attacks on Pre-trained Models
K. Kurita, P. Michel, and G. Neubig · 2004
Earlier work this paper cites.
Concealed Data Poisoning Attacks on NLP Models
E. Wallace, T. Z. Zhao, S. Feng, and S. Singh · 2010
Earlier work this paper cites.
Memory transformation and systems consolidation
G. Winocur and M. Moscovitch · 2011
Earlier work this paper cites.
Behavioral priming: It’s all in the mind, but whose mind?
S. Doyen · 2012
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories
M. Geva, R. Schuster, J. Berant, and O. Levy · 2012
Earlier work this paper cites.
Task-oriented intrinsic evaluation of semantic textual similarity
N. Reimers, P. Beyer, and I. Gurevych · 2016
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
O. Levy, M. Seo, E. Choi, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Locus coeruleus input to hippocampal ca3 drives single-trial learning of a novel context
A. Wagatsuma, T. Okuyama, C. Sun, L. M. Smith, K. Abe, and S. Tonegawa · 2018
Earlier work this paper cites.
Integration of new information in memory: new insights from a complementary learning systems perspective
J. L. McClelland, B. L. McNaughton, and A. K. Lampinen · 2019
Earlier work this paper cites.
Measuring and Improving Consistency in Pretrained Language Models
Y. Elazar, N. Kassner, S. Ravfogel, A. Ravichander, E. Hovy, H. Schütze, and Y. Goldberg · 2021
Earlier work this paper cites.
Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
M. Geva, A. Caciularu, K. R. Wang, and Y. Goldberg · 2022
Cited alongside, same era.
Biological underpinnings for lifelong learning machines
D. Kudithipudi, M. Aguilar-Simon, J. Babb, M. Bazhenov, D. Blackiston, J. Bongard, A. Brna, S. Chakravarthi Raja, N. Cheney, J. Clune, A. Daram, S. Fusi, P. Helfer, L. Kay, N. Ketz, Z. Kira, S. Kolouri, J. Krichmar, S. Kriegman, and H. Siegelmann · 2022
Cited alongside, same era.
Memory-Based Model Editing at Scale
E. Mitchell, C. Lin, A. Bosselut, C. D. Manning, and C. Finn · 2022
Cited alongside, same era.
Learning in deep neural networks and brains with similarity-weighted interleaved learning
R. Saxena, J. Shobe, and B. Mcnaughton · 2022
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. Canton Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. Singh Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Poisoning Language Models During Instruction Tuning
A. Wan, E. Wallace, S. Shen, and D. Klein · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2022
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
K. Ahn, X. Cheng, H. Daneshmand, and S. Sra · 2023
Cited alongside, same era.
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
Z. Allen-Zhu and Y. Li · 2023
Cited alongside, same era.
PaLM 2 Technical Report
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. El Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. Hernandez Abrego, J. Ahn, J. Austin, P. Barham, J. Botha, J. Bradbury, S. Brahma, K. Brooks, M. Catasta, Y. Cheng, C. Cherry, C. A. Choquette-Choo, A. Chowdhery, C. Crepy, S. Dave, M. Dehghani, S. Dev, J. Devlin, M. Díaz, N. Du, E. Dyer, V. Feinberg, F. Feng, V. Fienber, M. Freitag, X. Garcia, S. Gehrmann, L. Gonzalez, G. Gur-Ari, S. Hand, H. Hashemi, L. Hou, J. Howland, A. Hu, J. Hui, J. Hurwitz, M. Isard, A. Ittycheriah, M. Jagielski, W. Jia, K. Kenealy, M. Krikun, S. Kudugunta, C. Lan, K. Lee, B. Lee, E. Li, M. Li, W. Li, Y. Li, J. Li, H. Lim, H. Lin, Z. Liu, F. Liu, M. Maggioni, A. Mahendru, J. Maynez, V. Misra, M. Moussalem, Z. Nado, J. Nham, E. Ni, A. Nystrom, A. Parrish, M. Pellat, M. Polacek, A. Polozov, R. Pope, S. Qiao, E. Reif, B. Richter, P. Riley, A. Castro Ros, A. Roy, B. Saeta, R. Samuel, R. Shelby, A. Slone, D. Smilkov, D. R. So, D. Sohn, S. Tokumine, D. Valter, V. Vasudevan, K. Vodrahalli, X. Wang, P. Wang, Z. Wang, T. Wang, J. Wieting, Y. Wu, K. Xu, Y. Xu, L. Xue, P. Yin, J. Yu, Q. Zhang, S. Zheng, C. Zheng, W. Zhou, D. Zhou, S. Petrov, and Y. Wu · 2023
Cited alongside, same era.
Poisoning Web-Scale Training Datasets is Practical
N. Carlini, M. Jagielski, C. A. Choquette-Choo, D. Paleka, W. Pearce, H. Anderson, A. Terzis, K. Thomas, and F. Tramèr · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Gemini Team Google · 2023
Cited alongside, same era.
Dissecting Recall of Factual Associations in Auto-Regressive Language Models
M. Geva, J. Bastings, K. Filippova, and A. Globerson · 2023
Cited alongside, same era.
P. Hase, M. Bansal, B. Kim, and A. Ghandeharioun · 2023
Cited alongside, same era.
P. Yadav, D. Tam, L. Choshen, C. Raffel, and M. Bansal · 2023
Later among the works it cites.
Editing Large Language Models: Problems, Methods, and Opportunities
Y. Yao, P. Wang, B. Tian, S. Cheng, Z. Li, S. Deng, H. Chen, and N. Zhang · 2023
Later among the works it cites.
ALCUNA: Large Language Models Meet New Knowledge
X. Yin, B. Huang, and X. Wan · 2023
Later among the works it cites.
Detecting hallucinations in large language models using semantic entropy
S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal · 2024
Later among the works it cites.
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
Z. Gekhman, G. Yona, R. Aharoni, M. Eyal, A. Feder, R. Reichart, and J. Herzig · 2024
Later among the works it cites.
Gemma: Open Models Based on Gemini Research and Technology
Gemma Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, P. Tafti, L. Hussenot, P. G. Sessa, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. Héliou, A. Tacchetti, A. Bulanova, A. Paterson, B. Tsai, B. Shahriari, C. Le Lan, C. A. Choquette-Choo, C. Crepy, D. Cer, D. Ippolito, D. Reid, E. Buchatskaya, E. Ni, E. Noland, G. Yan, G. Tucker, G.-C. Muraru, G. Rozhdestvenskiy, H. Michalewski, I. Tenney, I. Grishchenko, J. Austin, J. Keeling, J. Labanowski, J.-B. Lespiau, J. Stanway, J. Brennan, J. Chen, J. Ferret, J. Chiu, J. Mao-Jones, K. Lee, K. Yu, K. Millican, L. Lowe Sjoesund, L. Lee, L. Dixon, M. Reid, M. Mikuła, M. Wirth, M. Sharman, N. Chinaev, N. Thain, O. Bachem, O. Chang, O. Wahltinez, P. Bailey, P. Michel, P. Yotov, R. Chaabouni, R. Comanescu, R. Jana, R. Anil, R. McIlroy, R. Liu, R. Mullins, S. L. Smith, S. Borgeaud, S. Girgin, S. Douglas, S. Pandya, S. Shakeri, S. De, T. Klimenko, T. Hennigan, V. Feinberg, W. Stokowiec, Y.-h. Chen, Z. Ahmed, Z. Gong, T. Warkentin, L. Peran, M. Giang, C. Farabet, O. Vinyals, J. Dean, K. Kavukcuoglu, D. Hassabis, Z. Ghahramani, D. Eck, J. Barral, F. Pereira, E. Collins, A. Joulin, N. Fiedel, E. Senter, A. Andreev, and K. Kenealy · 2024
Later among the works it cites.
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
A. Ghandeharioun, A. Caciularu, A. Pearce, L. Dixon, and M. Geva · 2024
Later among the works it cites.
Reverse Training to Nurse the Reversal Curse
O. Golovneva, Z. Allen-Zhu, J. Weston, and S. Sukhbaatar · 2024
Later among the works it cites.
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning
N. Mecklenburg, Y. Lin, X. Li, D. Holstein, L. Nunes, S. Malvar, B. Silva, R. Chandra, V. Aski, P. K. R. Yannam, T. Aktas, and T. Hendry · 2024
Later among the works it cites.
Continual Learning of Large Language Models: A Comprehensive Survey
H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y. Wang, Z. Wang, S. Ebrahimi, and H. Wang · 2024
Later among the works it cites.
Continual Learning for Large Language Models: A Survey
T. Wu, L. Luo, Y.-F. Li, S. Pan, T.-T. Vu, and G. Haffari · 2024
Later among the works it cites.
Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models
Y. Zhang, S. Li, J. Liu, P. Yu, Y. R. Fung, J. Li, M. Li, and H. Ji · 2024
Later among the works it cites.