Fetching the paper…
Reading the bibliography…
Latent Chain-of-Thought (Latent CoT) models promise efficient reasoning via continuous representations, yet exhibit puzzling performance inconsistencies: excelling at exploration (ProsQA: 97.0%) but failing at computation (GSM8K: 34.1%).
Improved algorithms for linear stochastic bandits
Abbasi-yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Deep learning and the information bottleneck principle, 2015
Tishby, N. and Zaslavsky, N · 2015
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Implicit chain of thought reasoning via knowledge distillation
Deng, Y., Prasad, K., Fernandez, R., Smolensky, P., Chaudhary, V., and Shieber, S · 2023
Earlier work this paper cites.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models
Hou, Y., Li, J., Fei, Y., Stolfo, A., Zhou, W., Zeng, G., Bosselut, A., and Sachan, M · 2023
Earlier work this paper cites.
A survey of reasoning with foundation models
Sun, J., Zheng, C., Xie, E., Liu, Z., Chu, R., Qiu, J., Xu, J., Ding, M., Li, H., Geng, M., et al · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models, 2023
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2023
Earlier work this paper cites.
Language models are hidden reasoners: Unlocking latent reasoning capabilities via self-rewarding
Chen, H., Feng, Y., Liu, Z., Yao, W., Prabhakar, A., Heinecke, S., Ho, R., Mui, P., Savarese, S., Xiong, C., et al · 2024
Earlier work this paper cites.
Compressed chain of thought: Efficient reasoning through dense representations
Cheng, J. and Van Durme, B · 2024
Earlier work this paper cites.
From explicit cot to implicit cot: Learning to internalize cot step by step
Deng, Y., Choi, Y., and Shieber, S · 2024
Earlier work this paper cites.
Think before you speak: Training language models with pause tokens
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2024
Earlier work this paper cites.
Thinking tokens for language modeling
Herel, D. and Mikolov, T · 2024
Earlier work this paper cites.
Expediting and elevating large language model reasoning via hidden chain-of-thought decoding
Liu, T., Chen, Z., Liu, Z., Tian, M., and Luo, W · 2024
Earlier work this paper cites.
Let’s think dot by dot: Hidden computation in transformer language models
Pfau, J., Merrill, W., and Bowman, S. R · 2024
Cited alongside, same era.
Guiding language model reasoning with planning tokens
Wang, X., Caccia, L., Ostapenko, O., Yuan, X., Wang, W. Y., and Sordoni, A · 2024
Cited alongside, same era.
Llava-o1: Let vision language models reason step-by-step
Xu, G., Jin, P., Hao, L., Song, Y., Sun, L., and Yuan, L · 2024
Cited alongside, same era.
Distilling system 2 into system 1
Yu, P., Xu, J., Weston, J. E., and Kulikov, I · 2024
Cited alongside, same era.
Quiet-STar: Language models can teach themselves to think before speaking
Zelikman, E., Harik, G. R., Shao, Y., Jayasiri, V., Haber, N., and Goodman, N · 2024
Cited alongside, same era.
Deepseek-v3.1 model card
DeepSeek-AI · 2025
CoTFormer: A chain of thought driven architecture with budget-adaptive computation cost at inference
Mohtashami, A., Pagliardini, M., and Jaggi, M · 2025
Later among the works it cites.
Introducing gpt-5
OpenAI · 2025
Later among the works it cites.
Beyond words: A latent memory approach to internal reasoning in llms
Orlicki, J. I · 2025
Later among the works it cites.
Reasoning to learn from latent thoughts
Ruan, Y., Band, N., Maddison, C. J., and Hashimoto, T · 2025
Later among the works it cites.
Token assorted: Mixing latent and text tokens for improved language model reasoning
Su, D., Zhu, H., Xu, Y., Jiao, J., Tian, Y., and Zheng, Q · 2025
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
When chain of thought is necessary, language models struggle to evade monitors
Emmons, S., Jenner, E., Elson, D. K., Saurous, R. A., Rajamanoharan, S., Chen, H., Shafkat, I., and Shah, R · 2025
Cited alongside, same era.
Continuous chain of thought enables parallel exploration and reasoning
Gozeten, H. A., Ildiz, M. E., Zhang, X., Harutyunyan, H., Rawat, A. S., and Oymak, S · 2025
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Cited alongside, same era.
Reconsidering overthinking: Penalizing internal and external redundancy in cot reasoning
Hong, J., Zhen, T., Chen, K., Liu, J., Zhu, W., Huo, J., Gao, Y., Wang, D., Wan, H., Yang, X., et al · 2025
Cited alongside, same era.
Scalable language models with posterior inference of latent thought vectors
Kong, D., Zhao, M., Xu, D., Pang, B., Wang, S., Honig, E., Si, Z., Li, C., Xie, J., Xie, S., et al · 2025
Cited alongside, same era.
A survey of personalized large language models: Progress and future directions
Liu, J., Qiu, Z., Li, Z., Dai, Q., Zhu, J., Hu, M., Yang, M., and King, I · 2025
Cited alongside, same era.
Sun, Y., Chen, Y., Li, Y., and Ding, B · 2025
Later among the works it cites.
Llm pretraining with continuous concepts
Tack, J., Lanchantin, J., Yu, J., Cohen, A., Kulikov, I., Lan, J., Hao, S., Tian, Y., Weston, J., and Li, X · 2025
Later among the works it cites.
Think silently, think fast: Dynamic latent compression of llm reasoning chains
Tan, W., Li, J., Ju, J., Luo, Z., Luan, J., and Song, R · 2025
Later among the works it cites.
How does transformer learn implicit reasoning?
Ye, J., Yao, Z., Huang, Z., Pan, L., Liu, J., Bai, Y., Xin, A., Weichuan, L., Che, X., Hou, L., et al · 2025
Later among the works it cites.
Enhancing auto-regressive chain-of-thought through loop-aligned reasoning
Yu, Q., He, Z., Li, S., Zhou, X., Zhang, J., Xu, J., and He, D · 2025
Later among the works it cites.
Don’t overthink it: A survey of efficient r1-style large reasoning models
Yue, L., Du, Y., Wang, Y., Gao, W., Yao, F., Wang, L., Liu, Y., Xu, Z., Liu, Q., Di, S., et al · 2025
Later among the works it cites.
Pretraining language models to ponder in continuous space
Zeng, B., Song, S., Huang, S., Wang, Y., Li, H., He, Z., Wang, X., Li, Z., and Lin, Z · 2025
Later among the works it cites.
Reasoning by superposition: A theoretical perspective on chain of continuous thought
Zhu, H., Hao, S., Hu, Z., Jiao, J., Russell, S., and Tian, Y · 2025
Later among the works it cites.