Fetching the paper…
Reading the bibliography…
As large language models (LLMs) continue to demonstrate remarkable abilities across various domains, computer scientists are developing methods to understand their cognitive processes, particularly concerning how (and if) LLMs internally represent their beliefs about the world.
Truth and probability
Ramsey, F. P. (1926) · 1926
Earlier work this paper cites.
La prévision: ses lois logiques, ses sources subjectives
De Finetti, B. (1937) · 1937
Earlier work this paper cites.
Probability, frequency and reasonable expectation
Cox, R. T. (1946) · 1946
Earlier work this paper cites.
An essentially complete class of admissible decision functions
Wald, A. (1947) · 1947
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Brier, G. W. (1950) · 1950
Earlier work this paper cites.
Mental events
Davidson, D. (1970) · 1970
Earlier work this paper cites.
The foundations of statistics
Savage, L. J. (1972) · 1972
Earlier work this paper cites.
Radical interpretation
Davidson, D. (1973) · 1973
Earlier work this paper cites.
On the very idea of a conceptual scheme
Davidson, D. (1974) · 1974
Earlier work this paper cites.
Radical interpretation
Lewis, D. (1974) · 1974
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Tversky, A. and D. Kahneman (1974) · 1974
Earlier work this paper cites.
A theory of the learnable
Valiant, L. G. (1984) · 1984
Earlier work this paper cites.
Consequentialist foundations for expected utility
Hammond, P. J. (1988) · 1988
Earlier work this paper cites.
Local vs. distributed coding
Thorpe, S. (1989) · 1989
Earlier work this paper cites.
The logic of decision
Jeffrey, R. C. (1990) · 1990
Earlier work this paper cites.
Clever bookies and coherent beliefs
Christensen, D. (1991) · 1991
Earlier work this paper cites.
A nonpragmatic vindication of probabilism
Joyce, J. M. (1998) · 1998
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N. (1999) · 1999
Earlier work this paper cites.
Measuring incoherence
Schervish, M. J., T. Seidenfeld, and J. B. Kadane (2002) · 2002
Earlier work this paper cites.
Compressed sensing
Donoho, D. L. (2006) · 2006
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Gneiting, T. and A. E. Raftery (2007) · 2007
Earlier work this paper cites.
Open problems in cooperative ai
Dafoe, A., E. Hughes, Y. Bachrach, T. Collins, K. R. McKee, J. Z. Leibo, K. Larson, and T. Graepel (2020) · 2012
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and S. Ben-David (2014) · 2014
Earlier work this paper cites.
Reasons without persons: Rationality, identity, and time
Hedden, B. (2015) · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Y. Bengio (2016) · 2016
Cited alongside, same era.
Accuracy and the Laws of Credence
Pettigrew, R. (2016) · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei (2017) · 2017
Cited alongside, same era.
The stability of belief: How rational belief coheres with probability
Leitgeb, H. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin (2017) · 2017
Cited alongside, same era.
Climbing towards nlu: On meaning, form, and understanding in the age of data
Bender, E. M. and A. Koller (2020) · 2020
The internal state of an llm knows when it’s lying
Azaria, A. and T. Mitchell (2023) · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah (2023) · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang (2023) · 2023
Later among the works it cites.
Localizing lying in llama: Understanding instructed dishonesty on true-false questions through prompting, probing, and patching
Campbell, J., R. Ren, and P. Guo (2023) · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsettled thoughts: A theory of degrees of rationality
Staffel, J. (2020) · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano (2020) · 2020
Cited alongside, same era.
Towards explainability for ai fairness
Zhou, J., F. Chen, and A. Holzinger (2020) · 2020
Cited alongside, same era.
Can language models encode perceptual structure without grounding? a case study in color
Abdou, M., A. Kulmizev, D. Hershcovich, S. Frank, E. Pavlick, and A. Søgaard (2021) · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., T. Gebru, A. McMillan-Major, and S. Shmitchell (2021) · 2021
Cited alongside, same era.
Truthful ai: Developing and governing ai that does not lie
Evans, O., O. Cotton-Barratt, L. Finnveden, A. Bales, A. Balwit, P. Wills, L. Righetti, and W. Saunders (2021) · 2021
Cited alongside, same era.
Cunningham, H., A. Ewart, L. Riggs, R. Huben, and L. Sharkey (2023) · 2023
Later among the works it cites.
Challenges with unsupervised llm knowledge discovery
Farquhar, S., V. Varma, Z. Kenton, J. Gasteiger, V. Mikulik, and R. Shah (2023) · 2023
Later among the works it cites.
The allure of simplicity: On interpretable machine learning models in healthcare
Grote, T. (2023) · 2023
Later among the works it cites.
Operationalising representation in natural language processing
Harding, J. (2023) · 2023
Later among the works it cites.
Probing the quantitative–qualitative divide in probabilistic reasoning
Ibeling, D., T. Icard, K. Mierzewski, and M. Mossé (2023) · 2023
Later among the works it cites.
Specific versus general principles for constitutional ai
Kundu, S., Y. Bai, S. Kadavath, A. Askell, A. Callahan, A. Chen, A. Goldie, A. Balwit, A. Mirhoseini, B. McLean, et al. (2023) · 2023
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K., A. K. Hopkins, D. Bau, F. Viégas, H. Pfister, and M. Wattenberg (2023) · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K., O. Patel, F. Viégas, H. Pfister, and M. Wattenberg (2023) · 2023
Later among the works it cites.
Mandelkern, M. and T. Linzen (2023) · 2023
Later among the works it cites.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Marks, S. and M. Tegmark (2023) · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Nanda, N., L. Chan, T. Lieberum, J. Smith, and J. Steinhardt (2023) · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Nanda, N., A. Lee, and M. Wattenberg (2023) · 2023
Later among the works it cites.
Symbols and grounding in large language models
Pavlick, E. (2023) · 2023
Later among the works it cites.
How to explain and justify almost any decision: Potential pitfalls for accountability in ai decision-making
Zhou, J. and T. Joachims (2023) · 2023
Later among the works it cites.
Still no lie detector for language models: Probing empirical and conceptual roadblocks
Levinstein, B. A. and D. A. Herrmann (2024) · 2024
Closest in time.
Ai deception: A survey of examples, risks, and potential solutions
Park, P. S., S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks (2024) · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan (2024) · 2024
Closest in time.
Honesty is the best policy: defining and mitigating ai deception
Ward, F., F. Toni, F. Belardinelli, and T. Everitt (2024) · 2024
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Zhong, Z., Z. Liu, M. Tegmark, and J. Andreas (2024) · 2024
Closest in time.