Fetching the paper…
Reading the bibliography…
This study explores how large language models (LLMs) encode interwoven scientific knowledge, using chemical elements and LLaMA-series models as a case study.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
S. Teerapittayanon, B. McDanel, and H.-T. Kung · 2016
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
S. Arora, Y. Li, Y. Liang, T. Ma, and A. Risteski · 2018
Earlier work this paper cites.
W.-T. Kao, T.-H. Wu, P.-H. Chi, C.-C. Hsieh, and H.-Y. Lee · 2020
Earlier work this paper cites.
Semantic memory: A review of methods, models, and current challenges
A. A. Kumar · 2021
Earlier work this paper cites.
Analyzing transformers in embedding space
G. Dar, M. Geva, A. Gupta, and J. Berant · 2022
Earlier work this paper cites.
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, et al · 2022
Earlier work this paper cites.
Engineering monosemanticity in toy models
A. S. Jermyn, N. Schiefer, and E. Hubinger · 2022
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
K. Li, A. K. Hopkins, D. Bau, F. Viégas, H. Pfister, and M. Wattenberg · 2022
Earlier work this paper cites.
Polysemanticity and capacity in neural networks
A. Scherlis, K. Sachan, A. S. Jermyn, J. Benton, and B. Shlegeris · 2022
Earlier work this paper cites.
Associative thinking at the core of creativity
R. E. Beaty and Y. N. Kenett · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Cited alongside, same era.
Language models represent space and time
W. Gurnee and M. Tegmark · 2023
Cited alongside, same era.
In-context learning creates task vectors
R. Hendel, M. Geva, and A. Globerson · 2023
Cited alongside, same era.
Linearity of relation decoding in transformer language models
E. Hernandez, A. S. Sharma, T. Haklay, K. Meng, M. Wattenberg, J. Andreas, Y. Belinkov, and D. Bau · 2023
A survey on evaluation of large language models
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al · 2024
Later among the works it cites.
The representation landscape of few-shot learning and fine-tuning in large language models
D. Doimo, A. Serra, A. Ansuini, and A. Cazzaniga · 2024
Later among the works it cites.
Layer skip: Enabling early exit inference and self-speculative decoding
M. Elhoushi, A. Shrivastava, D. Liskovich, B. Hosmer, B. Wasti, L. Lai, A. Mahmoud, B. Acun, S. Agarwal, A. Roman, et al · 2024
Later among the works it cites.
Not all language model features are linear
J. Engels, E. J. Michaud, I. Liao, W. Gurnee, and M. Tegmark · 2024
Later among the works it cites.
How large language models encode context knowledge? a layer-wise probing study
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ai alignment: A comprehensive survey
J. Ji, T. Qiu, B. Chen, B. Zhang, H. Lou, K. Wang, Y. Duan, Z. He, J. Zhou, Z. Zhang, et al · 2023
Cited alongside, same era.
Challenges and applications of large language models
J. Kaddour, J. Harris, M. Mozes, H. Bradley, R. Raileanu, and R. McHardy · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt · 2023
Cited alongside, same era.
Transformer circuit faithfulness metrics are not robust
N. Nanda et al · 2023
Cited alongside, same era.
A comprehensive overview of large language models
H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian · 2023
Cited alongside, same era.
Linear representations of sentiment in large language models
C. Tigges, O. J. Hollinsworth, A. Geiger, and N. Nanda · 2023
Cited alongside, same era.
Mechanistic interpretability for ai safety–a review
L. Bereska and E. Gavves · 2024
Cited alongside, same era.
T. Ju, W. Sun, W. Du, X. Yuan, Z. Ren, and G. Liu · 2024
Later among the works it cites.
Sorted llama: Unlocking the potential of intermediate layers of large language models for dynamic inference
P. Kavehzadeh, M. Valipour, M. Tahaei, A. Ghodsi, B. Chen, and M. Rezagholizadeh · 2024
Later among the works it cites.
Z. Liu, C. Kong, Y. Liu, and M. Sun · 2024
Later among the works it cites.
Rethinking interpretability in the era of large language models
C. Singh, J. P. Inala, M. Galley, R. Caruana, and J. Gao · 2024
Later among the works it cites.
Does representation matter? exploring intermediate layers in large language models
O. Skean, M. R. Arefin, Y. LeCun, and R. Shwartz-Ziv · 2024
Later among the works it cites.
Entropy law: The story behind data compression and llm performance
M. Yin, C. Wu, Y. Wang, H. Wang, W. Guo, Y. Wang, Y. Liu, R. Tang, D. Lian, and E. Chen · 2024
Later among the works it cites.
Investigating layer importance in large language models
Y. Zhang, Y. Dong, and K. Kawaguchi · 2024
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Z. Zhong, Z. Liu, M. Tegmark, and J. Andreas · 2024
Later among the works it cites.