Fetching the paper…
Reading the bibliography…
This study investigates the internal information flow of large language models (LLMs) while performing chain-of-thought (CoT) style reasoning.
A stochastic approximation method
Herbert E. Robbins. 1951 · 1951
Earlier work this paper cites.
PLS-regression: a basic tool of chemometrics
Svante Wold, Michael Sjöström, and Lennart Eriksson. 2001 · 2001
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Guillaume Alain and Yoshua Bengio. 2017a · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017b · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Probing neural network comprehension of natural language arguments
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Gender bias in contextualized word embeddings
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019 · 2019
Earlier work this paper cites.
interpreting GPT: the logit lens
nostalgebraist. 2020 · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L. Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, and 5 others. 2023 · 2023
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2023 · 2023
Earlier work this paper cites.
Localizing lying in Llama: Understanding instructed dishonesty on true-false questions through prompting, probing, and patching
James Campbell, Phillip Guo, and Richard Ren. 2023 · 2023
Earlier work this paper cites.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, and 1 others. 2023 · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023 · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023 · 2023
Cited alongside, same era.
Do Deep Neural Networks Capture Compositionality in Arithmetic Reasoning?
Keito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, and Kentaro Inui. 2023 · 2023
Cited alongside, same era.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023 · 2023
Monotonic representation of numeric attributes in language models
Benjamin Heinzerling and Kentaro Inui. 2024 · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on Gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca D. Dragan, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
Mistral NeMo
Mistral AI Team. 2024 · 2024
Closest in time.
Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024 · 2024
Closest in time.
Qwen2.5: A party of foundation models
Qwen Team. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness?
Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. 2023 · 2023
Cited alongside, same era.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan. 2023 · 2023
Cited alongside, same era.
Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023 · 2023
Cited alongside, same era.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman. 2023 · 2023
Cited alongside, same era.
Chain-of-thought unfaithfulness as disguised accuracy
Oliver Bentham, Nathan Stringham, and Ana Marasovic. 2024 · 2024
Cited alongside, same era.
Iteration head: A mechanistic study of chain-of-thought
Vivien Cabannes, Charles Arnal, Wassim Bouaziz, Xingyu Alice Yang, Francois Charton, and Julia Kempe. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, and 82 others. 2024 · 2024
Cited alongside, same era.
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, and 11 others. 2024 · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda. 2024 · 2024
Closest in time.
Knowing before saying: LLM representations encode information about chain-of-thought success before completion
Anum Afzal, Florian Matthes, Gal Chechik, and Yftah Ziser. 2025 · 2025
Closest in time.
Reasoning models don’t always say what they think
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vladimir Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, and Ethan Perez. 2025 · 2025
Closest in time.
Post-hoc reasoning in chain of thought
Kyle Cox. 2025 · 2025
Closest in time.
Analysing chain of thought dynamics: Active guidance or unfaithful post-hoc rationalisation?
Samuel Lewis-Lim, Xingwei Tan, Zhixue Zhao, and Nikolaos Aletras. 2025 · 2025
Closest in time.
Physics of language models: Part 2.1, grade-school math and the hidden reasoning process
Tian Ye, Zicheng Xu, Yuanzhi Li, and Zeyuan Allen-Zhu. 2025 · 2025
Closest in time.
Do LLMs really think step-by-step in implicit reasoning?
Yijiong Yu. 2025 · 2025
Closest in time.
Language models encode the value of numbers linearly
Fangwei Zhu, Damai Dai, and Zhifang Sui. 2025 · 2025
Closest in time.