Fetching the paper…
Reading the bibliography…
Modern AI models contain much of human knowledge, yet understanding of their internal representation of this knowledge remains elusive.
Learning with kernels , volume 4
A. J. Smola and B. Schölkopf · 1998
Earlier work this paper cites.
Scikit-learn: Machine learning in python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, et al · 2011
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
T. Mikolov, W.-t. Yih, and G. Zweig · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Z. Yun, Y. Chen, B. Olshausen, and Y. LeCun · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Y. Belinkov · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
C. Burns, H. Ye, D. Klein, and J. Steinhardt · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
The internal state of an LLM knows when it’s lying
A. Azaria and T. Mitchell · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
R. Huben, H. Cunningham, L. R. Smith, A. Ewart, and L. Sharkey · 2023
Cited alongside, same era.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
J. Li, X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen · 2023
Cited alongside, same era.
ToxicChat: Unveiling hidden challenges of toxicity detection in real-world user-AI conversation
Z. Lin, Z. Wang, Y. Tong, Y. Wang, Y. Guo, Y. Wang, and J. Shang · 2023
Cited alongside, same era.
Emergent linear representations in world models of self-supervised sequence models
N. Nanda, A. Lee, and M. Wattenberg · 2023
Cited alongside, same era.
OpenAI · 2024
Later among the works it cites.
Mechanism for feature learning in neural networks and backpropagation-free machine learning models
A. Radhakrishnan, D. Beaglehole, P. Pandit, and M. Belkin · 2024
Later among the works it cites.
Lynx: An open source hallucination evaluation model
S. S. Ravi, B. Mielczarek, A. Kannappan, D. Kiela, and R. Qian · 2024
Later among the works it cites.
Steering llama 2 via contrastive activation addition
N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. Turner · 2024
Later among the works it cites.
Llm-check: Investigating detection of hallucinations in large language models
G. Sriramanan, S. Bharti, V. S. Sadasivan, S. Saha, P. Kattakinda, and S. Feizi · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Activation addition: Steering language models without optimization
A. M. Turner, L. Thiergart, G. Leech, D. Udell, J. J. Vazquez, U. Mini, and M. MacDiarmid · 2023
Cited alongside, same era.
AI@Meta · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Cited alongside, same era.
Estimating knowledge in large language models without generating a single token
D. Gottesman and M. Geva · 2024
Cited alongside, same era.
Sparse crosscoders for cross-layer features and model diffing
J. Lindsey, A. Templeton, J. Marcus, T. Conerly, J. Batson, and C. Olah · 2024
Cited alongside, same era.
Fine-grained hallucination detection and editing for language models
A. Mishra, A. Asai, V. Balachandran, Y. Wang, G. Neubig, Y. Tsvetkov, and H. Hajishirzi · 2024
Cited alongside, same era.
RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models
C. Niu, Y. Wu, J. Zhu, S. Xu, K. Shum, R. Zhong, J. Song, and T. Zhang · 2024
Cited alongside, same era.
Z. Zhu, Y. Yang, and Z. Sun · 2024
Later among the works it cites.
Detecting strategic deception using linear probes
N. Goldowsky-Dill, B. Chughtai, S. Heimersheim, and M. Hobbhahn · 2025
Closest in time.
Hackerrank coding challenges
HackerRank · 2025
Closest in time.
Negative results for sparse autoencoders on downstream tasks (and deprioritising sae research)
L. Smith, S. Rajamanoharan, A. Conmy, C. McDougall, J. Kramar, T. Lieberum, R. Shah, and N. Nanda · 2025
Closest in time.
Improving instruction-following in language models through activation steering
A. Stolfo, V. Balachandran, S. Yousefi, E. Horvitz, and B. Nushi · 2025
Closest in time.
Axbench: Steering LLMs? Even simple baselines outperform sparse autoencoders
Z. Wu, A. Arora, A. Geiger, Z. Wang, J. Huang, D. Jurafsky, C. D. Manning, and C. Potts · 2025
Closest in time.