Fetching the paper…
Reading the bibliography…
Pre-trained language models (PLMs) contain vast amounts of factual knowledge, but how the knowledge is stored in the parameters remains unclear.
Language Models as Knowledge Bases?
Petroni, F.; Rocktäschel, T.; Lewis, P.; Bakhtin, A.; Wu, Y.; Miller, A. H.; and Riedel, S. 2019b · 1909
Earlier work this paper cites.
How Can We Know What Language Models Know?
Jiang, Z.; Xu, F. F.; Araki, J.; and Neubig, G. 2020 · 1911
Earlier work this paper cites.
Measures of degeneracy and redundancy in biological networks
Tononi, G.; Sporns, O.; and Edelman, G. M. 1999 · 1999
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories
Geva, M.; Schuster, R.; Berant, J.; and Levy, O. 2021 · 2012
Earlier work this paper cites.
Degeneracy: Demystifying and destigmatizing a core concept in systems biology
Mason, P. H. 2015 · 2015
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Sundararajan, M.; Taly, A.; and Yan, Q. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Gradient-based attribution methods
Ancona, M.; et al. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?
Petroni, F.; Rocktäschel, T.; Riedel, S.; Lewis, P.; Bakhtin, A.; Wu, Y.; and Miller, A. 2019a · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Earlier work this paper cites.
How Context Affects Language Models’ Factual Predictions
Petroni, F.; Lewis, P.; Piktus, A.; Rocktäschel, T.; Wu, Y.; Miller, A. H.; and Riedel, S. 2020 · 2020
Earlier work this paper cites.
On Negative Interference in Multilingual Models: Findings and A Meta-Learning Treatment
Wang, Z.; Lipton, Z. C.; and Tsvetkov, Y. 2020 · 2020
Earlier work this paper cites.
Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models
Kassner, N.; Dufter, P.; and Schütze, H. 2021 · 2021
Earlier work this paper cites.
Discretized Integrated Gradients for Explaining Language Models
Sanyal, S.; and Ren, X. 2021 · 2021
Cited alongside, same era.
Language Models as Agent Models
Andreas, J. 2022 · 2022
Cited alongside, same era.
Knowledge Neurons in Pretrained Transformers
Dai, D.; Dong, L.; Hao, Y.; Sui, Z.; Chang, B.; and Wei, F. 2022 · 2022
Cited alongside, same era.
OpenAI invites everyone to test ChatGPT, a new AI-powered chatbot—with amusing results
Edwards, B. 2022 · 2022
Cited alongside, same era.
Why large language models like ChatGPT are bullshit artists
Lakshmanan, L. 2022 · 2022
Cited alongside, same era.
How Pre-trained Language Models Capture Factual Knowledge? A Causal-Inspired Analysis
A survey on knowledge-enhanced pre-trained language models
Zhen, C.; Shang, Y.; Liu, X.; Li, Y.; Chen, Y.; and Zhang, D. 2022 · 2022
Later among the works it cites.
The Life Cycle of Knowledge in Big Language Models: A Survey
Cao, B.; Lin, H.; Han, X.; and Sun, L. 2023 · 2023
Closest in time.
Why ChatGPT and Bing Chat are so good at making things up
Edwards, B. 2023 · 2023
Closest in time.
Sequential Integrated Gradients: a simple but effective method for explaining language models
Enguehard, J. 2023 · 2023
Closest in time.
Hase, P.; Bansal, M.; Kim, B.; and Ghandeharioun, A. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, S.; Li, X.; Shang, L.; Dong, Z.; Sun, C.; Liu, B.; Ji, Z.; Jiang, X.; and Liu, Q. 2022 · 2022
Cited alongside, same era.
The Effective coalitions of Shapley value For Integrated Gradients
Liu, S.; Fan, C.; Xiong, Y.; Wang, M.; Hu, Y.; Lv, T.; Chen, Z.; Wu, R.; and Gao, Y. 2022 · 2022
Cited alongside, same era.
A rigorous study of integrated gradients method and extensions to internal neuron attributions
Lundstrom, D. D.; Huang, T.; and Razaviyayn, M. 2022 · 2022
Cited alongside, same era.
The new chatbots could change the world. Can you trust them
Metz, C. 2022 · 2022
Cited alongside, same era.
Fast Model Editing at Scale
Mitchell, E.; Lin, C.; Bosselut, A.; Finn, C.; and Manning, C. D. 2022 · 2022
Cited alongside, same era.
Google vs. ChatGPT: Here’s what happened when I swapped services for a day
Pitt, S. 2022 · 2022
Cited alongside, same era.
mGPT: Few-Shot Learners Go Multilingual
Shliazhko, O.; Fenogenova, A.; Tikhonova, M.; Mikhailov, V.; Kozlova, A.; and Shavrina, T. 2022 · 2022
Cited alongside, same era.
Closest in time.
Large language models struggle to learn long-tail knowledge
Kandpal, N.; Deng, H.; Roberts, A.; Wallace, E.; and Raffel, C. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023 · 2023
Closest in time.
Scientific Fact-Checking: A Survey of Resources and Approaches
Vladika, J.; and Matthes, F. 2023 · 2023
Closest in time.
Language Anisotropic Cross-Lingual Model Editing
Xu, Y.; Hou, Y.; Che, W.; and Zhang, M. 2023 · 2023
Closest in time.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Zhou, C.; Li, Q.; Li, C.; Yu, J.; Liu, Y.; Wang, G.; Zhang, K.; Ji, C.; Yan, Q.; He, L.; et al. 2023 · 2023
Closest in time.