Fetching the paper…
Reading the bibliography…
There has been recent interest in whether large language models (LLMs) can introspect about their own internal states.
A history of introspection
Edwin G Boring · 1953
Earlier work this paper cites.
Syntactic Structures
N. Chomsky · 1957
Earlier work this paper cites.
Telling more than we can know: Verbal reports on mental processes
Richard E. Nisbett and Timothy D. Wilson · 1977
Earlier work this paper cites.
Brown corpus manual
W Nelson Francis and Henry Kucera · 1979
Earlier work this paper cites.
Behaviorism and the mind: A (limited) call for a return to introspection
David A Lieberman · 1979
Earlier work this paper cites.
Representations and misrepresentations
BF Skinner · 1984
Earlier work this paper cites.
Introspection
Alex Byrne · 2005
Earlier work this paper cites.
The unreliability of naive introspection
Eric Schwitzgebel · 2008
Earlier work this paper cites.
The need for quantitative methods in syntax and semantics research
E. Gibson and E. Fedorenko · 2010
Earlier work this paper cites.
A validation of Amazon Mechanical Turk for the collection of acceptability judgments in linguistic theory
Jon Sprouse · 2011
Earlier work this paper cites.
A comparison of informal and formal acceptability judgments using a random sample from linguistic inquiry 2001–2010
Jon Sprouse, Carson T Schütze, and Diogo Almeida · 2013
Earlier work this paper cites.
Nisbett and Wilson (1977) Revisited: The Little That We Can Know and Can Tell
Christopher C. Berger, Tara C. Dennehy, John A. Bargh, and Ezequiel Morsella · 2016
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg · 2016
Earlier work this paper cites.
Snap judgments: A small n acceptability paradigm (snap) for linguistic acceptability judgments
Kyle Mahowald, Peter Graff, Jeremy Hartman, and Edward Gibson · 2016
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen · 2018
Earlier work this paper cites.
Introspection as a methodology in linguistics
Leonard Talmy · 2018
Earlier work this paper cites.
What do rnn language models learn about filler–gap dependencies?
Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell · 2018
Cited alongside, same era.
Neural language models as psycholinguistic subjects: Representations of syntactic state
Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy · 2019
Cited alongside, same era.
A survey of the state of explainable ai for natural language processing
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen · 2020
Cited alongside, same era.
A Systematic Assessment of Syntactic Generalization in Neural Language Models
Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy · 2020
Cited alongside, same era.
BLiMP: The Benchmark of Linguistic Minimal Pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman · 2020
Cited alongside, same era.
Llama 3 Model Card
AI@Meta · 2024
Later among the works it cites.
Auxiliary task demands mask the capabilities of smaller language models
Jennifer Hu and Michael Frank · 2024
Later among the works it cites.
Language models align with human judgments on key grammatical constructions
Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova, and Roger Levy · 2024
Later among the works it cites.
Benchmarking Cognitive Biases in Large Language Models as Evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang · 2024
Later among the works it cites.
Reply to Hu et al.: Applying different evaluation standards to humans vs. Large Language Models overestimates AI performance
Evelina Leivada, Fritz Günther, and Vittoria Dentella · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Cited alongside, same era.
minicons: Enabling flexible behavioral and representational analyses of transformer language models
Kanishka Misra · 2022
Cited alongside, same era.
Systematic testing of three language models reveals low language accuracy, absence of response stability, and a yes-response bias
Vittoria Dentella, Fritz Günther, and Evelina Leivada · 2023
Cited alongside, same era.
Explainable ai (xai): Core ideas, techniques, and solutions
Rudresh Dwivedi, Devam Dave, Het Naik, Smiti Singhal, Rana Omer, Pankesh Patel, Bin Qian, Zhenyu Wen, Tejal Shah, Graham Morgan, et al · 2023
Cited alongside, same era.
Prompting is not a substitute for probability measurements in large language models
Jennifer Hu and Roger Levy · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Why Large Language Models Are Poor Theories of Human Linguistic Cognition: A Reply to Piantadosi
Roni Katzir · 2023
Cited alongside, same era.
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, et al · 2024
Later among the works it cites.
LLM Evaluators Recognize and Favor Their Own Generations
Arjun Panickssery, Samuel R. Bowman, and Shi Feng · 2024
Later among the works it cites.
Introspection
Eric Schwitzgebel · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
I am a Strange Dataset: Metalinguistic Tests for Language Models
Tristan Thrush, Jared Moore, Miguel Monares, Christopher Potts, and Douwe Kiela · 2024
Later among the works it cites.
“my answer is C”: First-token probabilities do not match text answers in instruction-tuned language models
Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank · 2024
Later among the works it cites.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zhihao Fan · 2024
Later among the works it cites.
Which of these best describes multiple choice evaluation with LLMs? a) forced B) flawed C) fixable D) all of the above
Nishant Balepur, Rachel Rudinger, and Jordan Lee Boyd-Graber · 2025
Closest in time.
Tell me about yourself: LLMs are aware of their learned behaviors
Jan Betley, Xuchan Bao, Martín Soto, Anna Sztyber-Betley, James Chua, and Owain Evans · 2025
Closest in time.
Looking inward: Language models can learn about themselves by introspection
Felix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans · 2025
Closest in time.
How linguistics learned to stop worrying and love the language models
Richard Futrell and Kyle Mahowald · 2025
Closest in time.