Fetching the paper…
Reading the bibliography…
We consider the questions of whether or not large language models (LLMs) have beliefs, and, if they do, how we might measure them.
Risks from learned optimization in advanced machine learning systems
Hubinger, E., C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant (2019) · 1906
Earlier work this paper cites.
The theoretician’s dilemma: A study in the logic of theory construction
Hempel, C. G. (1958) · 1958
Earlier work this paper cites.
Word and object
Quine, W. V. O. (1960) · 1960
Earlier work this paper cites.
Natural kinds
Quine, W. V. (1969) · 1969
Earlier work this paper cites.
The foundations of statistics
Savage, L. J. (1972) · 1972
Earlier work this paper cites.
The framing of decisions and the psychology of choice
Tversky, A. and D. Kahneman (1981) · 1981
Earlier work this paper cites.
Reality and representation
Papineau, D. (1988) · 1988
Earlier work this paper cites.
The logic of decision
Jeffrey, R. C. (1990) · 1990
Earlier work this paper cites.
The fragmentation of reason: Preface to a pragmatic theory of cognitive evaluation
Stich, S. P. (1990) · 1990
Earlier work this paper cites.
Signal, decision, action
Godfrey-Smith, P. (1991) · 1991
Earlier work this paper cites.
The adaptive advantage of learning and a priori prejudice
Sober, E. (1994) · 1994
Earlier work this paper cites.
White queen psychology and other essays for Alice
Millikan, R. G. (1995) · 1995
Earlier work this paper cites.
Complexity and the Function of Mind in Nature
Godfrey-Smith, P. (1998) · 1998
Earlier work this paper cites.
When is it selectively advantageous to have true beliefs? sandwiching the better safe than sorry argument
Stephens, C. L. (2001) · 2001
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Gneiting, T. and A. E. Raftery (2007) · 2007
Earlier work this paper cites.
Social interaction and the evolution of learning rules
Smead, R. S. (2009) · 2009
Cited alongside, same era.
Evolution and the normativity of epistemic reasons
Street, S. (2009) · 2009
Cited alongside, same era.
Learning word vectors for sentiment analysis
Maas, A., R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts (2011) · 2011
Cited alongside, same era.
An introduction to latent variable models
Everett, B. (2013) · 2013
Cited alongside, same era.
In defence of instrumentalism about epistemic normativity
Cowie, C. (2014) · 2014
Cited alongside, same era.
The role of social interaction in the evolution of learning
Smead, R. (2015) · 2015
Cited alongside, same era.
Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources
Lieder, F. and T. L. Griffiths (2020) · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., T. Gebru, A. McMillan-Major, and S. Shmitchell (2021) · 2021
Later among the works it cites.
Arc’s first technical report: Eliciting latent knowledge
Christiano, P., M. Xu, and A. Cotra (2021) · 2021
Later among the works it cites.
Truthful ai: Developing and governing ai that does not lie
Evans, O., O. Cotton-Barratt, L. Finnveden, A. Bales, A. Balwit, P. Wills, L. Righetti, and W. Saunders (2021) · 2021
Later among the works it cites.
Gpt-j-6b: A 6 billion parameter autoregressive language model
Wang, B. and A. Komatsuzaki (2021, May) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alain, G. and Y. Bengio (2016) · 2016
Cited alongside, same era.
Deep learning
Goodfellow, I., Y. Bengio, and A. Courville (2016) · 2016
Cited alongside, same era.
Truth and probability
Ramsey, F. P. (2016) · 2016
Cited alongside, same era.
Attention is all you need
Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin (2017) · 2017
Cited alongside, same era.
Recognition in terra incognita
Beery, S., G. van Horn, and P. Perona (2018) · 2018
Cited alongside, same era.
Ten great ideas about chance
Diaconis, P. and B. Skyrms (2018) · 2018
Cited alongside, same era.
Xie, S. M., A. Raghunathan, P. Liang, and T. Ma (2021) · 2021
Later among the works it cites.
Discovering latent knowledge in language models without supervision
Burns, C., H. Ye, D. Klein, and J. Steinhardt (2022) · 2022
Later among the works it cites.
Talking about large language models
Shanahan, M. (2022) · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer (2022) · 2022
Later among the works it cites.
The internal state of an llm knows when its lying
Azaria, A. and T. Mitchell (2023) · 2023
Closest in time.
Operationalising representation in natural language processing
Harding, J. (2023) · 2023
Closest in time.
Survey of hallucination in natural language generation
Ji, Z., N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung (2023) · 2023
Closest in time.
A latent space theory for emergent abilities in large language models
Jiang, H. (2023) · 2023
Closest in time.
A conceptual guide to transformers
Levinstein, B. (2023) · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample (2023) · 2023
Closest in time.