Fetching the paper…
Reading the bibliography…
Latent Chain-of-Thought (Latent-CoT) aims to enable step-by-step computation without emitting long rationales, yet its mechanisms remain unclear.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
interpreting GPT: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Earlier work this paper cites.
Transformerlens
Nanda, N. and Bloom, J · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models
Hou, Y., Li, J., Fei, Y., Stolfo, A., Zhou, W., Zeng, G., Bosselut, A., and Sachan, M · 2023
Earlier work this paper cites.
Dissecting chain-of-thought: Compositionality through in-context filtering and learning
Li, Y., Sreenivasan, K., Giannou, A., Papailiopoulos, D., and Oymak, S · 2023
Earlier work this paper cites.
The expressive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2023
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., et al · 2023
Earlier work this paper cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Cited alongside, same era.
Hopping too late: Exploring the limitations of large language models on multi-hop queries
Biran, E., Gottesman, D., Yang, S., Geva, M., and Globerson, A · 2024
Cited alongside, same era.
A mechanistic analysis of a transformer trained on a symbolic multi-step reasoning task
Brinkmann, J., Sheshadri, A., Levoso, V., Swoboda, P., and Bartelt, C · 2024
Cited alongside, same era.
Iteration head: A mechanistic study of chain-of-thought
Cabannes, V., Arnal, C., Bouaziz, W., Yang, X., Charton, F., and Kempe, J · 2024
Cited alongside, same era.
From explicit cot to implicit cot: Learning to internalize cot step by step
Rai, D. and Yao, Z · 2024
Later among the works it cites.
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning
Sprague, Z., Yin, F., Rodriguez, J. D., Jiang, D., Wadhwa, M., Singhal, P., Zhao, X., Ye, X., Mahowald, K., and Durrett, G · 2024
Later among the works it cites.
Do llms really think step-by-step in implicit reasoning?
Yu, Y · 2024
Later among the works it cites.
Uncovering latent chain of thought vectors in language models
Zhang, J. and Viteri, S · 2024
Later among the works it cites.
Chain-of-thought reasoning in the wild is not always faithful
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deng, Y., Choi, Y., and Shieber, S · 2024
Cited alongside, same era.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
Dutta, S., Singh, J., Chakrabarti, S., and Chakraborty, T · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Cited alongside, same era.
Investigating multi-hop factual shortcuts in knowledge editing of large language models
Ju, T., Chen, Y., Yuan, X., Zhang, Z., Du, W., Zheng, Y., and Liu, G · 2024
Cited alongside, same era.
Think-to-talk or talk-to-think? when llms come up with an answer in multi-step reasoning
Kudo, K., Aoki, Y., Kuribayashi, T., Sone, S., Taniguchi, M., Brassard, A., Sakaguchi, K., and Inui, K · 2024
Cited alongside, same era.
Understanding and patching compositional reasoning in llms
Li, Z., Jiang, G., Xie, H., Song, L., Lian, D., and Wei, Y · 2024
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al
Cited in the paper.
Arcuschin, I., Janiak, J., Krzyzanowski, R., Rajamanoharan, S., Nanda, N., and Conmy, A · 2025
Later among the works it cites.
Reasoning with latent thoughts: On the power of looped transformers
Saunshi, N., Dikkala, N., Li, Z., Kumar, S., and Reddi, S. J · 2025
Later among the works it cites.
Codi: Compressing chain-of-thought into continuous space via self-distillation
Shen, Z., Yan, H., Zhang, L., Hu, Z., Du, Y., and He, Y · 2025
Later among the works it cites.
Do large language models perform latent multi-hop reasoning without exploiting shortcuts?
Yang, S., Kassner, N., Gribovskaya, E., Riedel, S., and Geva, M · 2025
Later among the works it cites.
Internal chain-of-thought: Empirical evidence for layer-wise subtask scheduling in llms
Yang, Z., Li, J., Xia, S., and Hu, X · 2025
Later among the works it cites.
How do transformers learn implicit reasoning?
Ye, J., Yao, Z., Huang, Z., Pan, L., Liu, J., Bai, Y., Xin, A., Weichuan, L., Che, X., Hou, L., et al · 2025
Later among the works it cites.
Back attention: Understanding and enhancing multi-hop reasoning in large language models
Yu, Z., Belinkov, Y., and Ananiadou, S · 2025
Later among the works it cites.