Fetching the paper…
Reading the bibliography…
Chain-of-thought (CoT) prompting boosts Large Language Models accuracy on multi-step tasks, yet whether the generated "thoughts" reflect the true internal reasoning process is unresolved.
Training Verifiers to Solve Math Word Problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021 · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Geiger, A.; Lu, H.; Icard, T.; and Potts, C. 2021 · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y. 2022 · 2022
Earlier work this paper cites.
Elhage, N.; Hume, T.; Olsson, C.; Schiefer, N.; Henighan, T.; Kravec, S.; Hatfield-Dodds, Z.; Lasenby, R.; Drain, D.; Chen, C.; et al. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Meng, K.; Bau, D.; Andonian, A.; and Belinkov, Y. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Faithfulness tests for natural language explanations
Atanasova, P.; Camburu, O.-M.; Lioma, C.; Lukasiewicz, T.; Simonsen, J. G.; and Augenstein, I. 2023 · 2023
Earlier work this paper cites.
Language models can explain neurons in language models
Bills, S.; Cammarata, N.; Mossing, D.; Tillman, H.; Gao, L.; Goh, G.; Sutskever, I.; Leike, J.; Wu, J.; and Saunders, W. 2023 · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; Askell, A.; et al. 2023 · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H.; Ewart, A.; Riggs, L.; Huben, R.; and Sharkey, L. 2023 · 2023
Earlier work this paper cites.
Causal abstraction: A theoretical foundation for mechanistic interpretability
Geiger, A.; Ibeling, D.; Zur, A.; Chaudhary, M.; Chauhan, S.; Huang, J.; Arora, A.; Wu, Z.; Goodman, N.; Potts, C.; et al. 2023 · 2023
Earlier work this paper cites.
Localizing model behavior with path patching
Goldowsky-Dill, N.; MacLeod, C.; Sato, L.; and Arora, A. 2023 · 2023
Earlier work this paper cites.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Hanna, M.; Liu, O.; and Variengien, A. 2023 · 2023
Earlier work this paper cites.
Makelov, A.; Lange, G.; and Nanda, N. 2023 · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability, 2023
Nanda, N.; Chan, L.; Lieberum, T.; Smith, J.; and Steinhardt, J. 2023 · 2023
Cited alongside, same era.
A reply to makelov et al.(2023)’s” interpretability illusion” arguments
Wu, Z.; Geiger, A.; Huang, J.; Arora, A.; Icard, T.; Potts, C.; and Goodman, N. D. 2024 · 2023
Cited alongside, same era.
Towards best practices of activation patching in language models: Metrics and methods
Zhang, F.; and Nanda, N. 2023 · 2023
Cited alongside, same era.
Faithfulness vs. plausibility: On the (un) reliability of explanations from large language models
Leveraging LLMs for Hypothetical Deduction in Logical Inference: A Neuro-Symbolic Approach
Li, Q.; Li, J.; Liu, T.; Zeng, Y.; Cheng, M.; Huang, W.; and Liu, Q. 2024 · 2024
Later among the works it cites.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Marks, S.; Rager, C.; Michaud, E. J.; Belinkov, Y.; Bau, D.; and Mueller, A. 2024 · 2024
Later among the works it cites.
Analyzing (In) Abilities of SAEs via Formal Languages
Menon, A.; Shrivastava, M.; Krueger, D.; and Lubana, E. S. 2024 · 2024
Later among the works it cites.
Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning
Paul, D.; West, R.; Bosselut, A.; and Faltings, B. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Agarwal, C.; Tanneru, S. H.; and Lakkaraju, H. 2024 · 2024
Cited alongside, same era.
Mechanistic Interpretability for AI Safety–A Review
Bereska, L.; and Gavves, E. 2024 · 2024
Cited alongside, same era.
Identifying functionally important features with end-to-end sparse dictionary learning
Braun, D.; Taylor, J.; Goldowsky-Dill, N.; and Sharkey, L. 2024 · 2024
Cited alongside, same era.
FaithLM: Towards faithful explanations for large language models
Chuang, Y.-N.; Wang, G.; Chang, C.-Y.; Tang, R.; Zhong, S.; Yang, F.; Du, M.; Cai, X.; and Hu, X. 2024 · 2024
Cited alongside, same era.
Sparse autoencoders reveal temporal difference learning in large language models
Demircan, C.; Saanum, T.; Jagadish, A. K.; Binz, M.; and Schulz, E. 2024 · 2024
Cited alongside, same era.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
Dutta, S.; Singh, J.; Chakrabarti, S.; and Chakraborty, T. 2024 · 2024
Cited alongside, same era.
Finding alignments between interpretable causal variables and distributed neural representations
Geiger, A.; Wu, Z.; Potts, C.; Icard, T.; and Goodman, N. 2024 · 2024
Cited alongside, same era.
How to use and interpret activation patching
Heimersheim, S.; and Nanda, N. 2024 · 2024
Cited alongside, same era.
Plaat, A.; Wong, A.; Verberne, S.; Broekens, J.; van Stein, N.; and Back, T. 2024 · 2024
Later among the works it cites.
Siegel, N. Y.; Camburu, O.-M.; Heess, N.; and Perez-Ortiz, M. 2024 · 2024
Later among the works it cites.
Probing Language Models on Their Knowledge Source
Tighidet, Z.; Mogini, A.; Mei, J.; Piwowarski, B.; and Gallinari, P. 2024 · 2024
Later among the works it cites.
Faithful logical reasoning via symbolic chain-of-thought
Xu, J.; Fei, H.; Pan, L.; Liu, Q.; Lee, M.-L.; and Hsu, W. 2024 · 2024
Later among the works it cites.
Dissociation of faithful and unfaithful reasoning in llms
Yee, E.; Li, A.; Tang, C.; Jung, Y. H.; Paturi, R.; and Bergen, L. 2024 · 2024
Later among the works it cites.
Yeo, W. J.; Satapathy, R.; and Cambria, E. 2024 · 2024
Later among the works it cites.
Tokenized SAEs: Disentangling SAE Reconstructions
Dooms, T.; and Wilhelm, D. 2025 · 2025
Closest in time.
Walk the talk? Measuring the faithfulness of large language model explanations
Matton, K.; Ness, R. O.; Guttag, J.; and Kıcıman, E. 2025 · 2025
Closest in time.
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
Shojaee, P.; Mirzadeh, I.; Alizadeh, K.; Horton, M.; Bengio, S.; and Farajtabar, M. 2025 · 2025
Closest in time.