Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) excel at reasoning and planning when trained on chainof-thought (CoT) data, where the step-by-step thought process is explicitly outlined by text tokens.
A formal basis for the heuristic determination of minimum cost paths
Hart, P. E., Nilsson, N. J., and Raphael, B · 1968
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
Saxton, D., Grefenstette, E., Hill, F., and Kohli, P · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Efficient planning in a compact latent action space
Jiang, Z., Zhang, T., Janner, M., Li, Y., Rocktäschel, T., Grefenstette, E., and Tian, Y · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Saparov, A. and He, H · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Llemma: An open language model for mathematics
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2023
Earlier work this paper cites.
Theoremqa: A theorem-driven question answering dataset
Chen, W., Yin, M., Ku, M., Lu, P., Wan, Y., Ma, X., Xu, J., Wang, X., and Xia, T · 2023
Earlier work this paper cites.
Implicit chain of thought reasoning via knowledge distillation
Deng, Y., Prasad, K., Fernandez, R., Smolensky, P., Chaudhary, V., and Shieber, S · 2023
Earlier work this paper cites.
Think before you speak: Training language models with pause tokens
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2023
Earlier work this paper cites.
H-gap: Humanoid control with a generalist planner
Jiang, Z., Xu, Y., Wagener, N., Luo, Y., Janner, M., Grefenstette, E., Rocktäschel, T., and Tian, Y · 2023
Earlier work this paper cites.
Kim, S., Joo, S. J., Kim, D., Jang, J., Ye, S., Shin, J., and Seo, M · 2023
Earlier work this paper cites.
Guiding language model reasoning with planning tokens
Wang, X., Caccia, L., Ostapenko, O., Yuan, X., Wang, W. Y., and Sordoni, A · 2023
Cited alongside, same era.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2023
Cited alongside, same era.
Mammoth: Building math generalist models through hybrid instruction tuning
Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Cited alongside, same era.
Large concept models: Language modeling in a sentence representation space
Barrault, L., Duquenne, P.-A., Elbayad, M., Kozhevnikov, A., Alastruey, B., Andrews, P., Coria, M., Couairon, G., Costa-jussà, M. R., Dale, D., et al · 2024
Cited alongside, same era.
Scaling instruction-finetuned language models
Zebralogic: Benchmarking the logical reasoning ability of language models, 2024
Lin, B. Y., Bras, R. L., and Choi, Y · 2024
Later among the works it cites.
Deliberation in latent space via differentiable cache augmentation
Liu, L., Pfeiffer, J., Wu, J., Xie, J., and Szlam, A · 2024
Later among the works it cites.
Finemath: the finest collection of mathematical content, 2024
Lozhkov, A., Ben Allal, L., Bakouch, E., von Werra, L., and Wolf, T · 2024
Later among the works it cites.
Byte latent transformer: Patches scale better than tokens
Pagnoni, A., Pasunuru, R., Rodriguez, P., Nguyen, J., Muller, B., Li, M., Zhou, C., Yu, L., Weston, J., Zettlemoyer, L., Ghosh, G., Lewis, M., Holtzman, A., and Iyer, S · 2024
Later among the works it cites.
Let’s think dot by dot: Hidden computation in transformer language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Cited alongside, same era.
From explicit cot to implicit cot: Learning to internalize cot step by step
Deng, Y., Choi, Y., and Shieber, S · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., and Choi, Y · 2024
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Feng, G., Zhang, B., Gu, Y., Ye, H., He, D., and Wang, L · 2024
Cited alongside, same era.
Stream of search (sos): Learning to search in language
Gandhi, K., Lee, D., Grand, G., Liu, M., Cheng, W., Sharma, A., and Goodman, N. D · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Cited alongside, same era.
He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., et al · 2024
Cited alongside, same era.
Pfau, J., Merrill, W., and Bowman, S. R · 2024
Later among the works it cites.
Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces
Su, D., Sukhbaatar, S., Rabbat, M., Tian, Y., and Zheng, Q · 2024
Later among the works it cites.
Mathscale: Scaling instruction tuning for mathematical reasoning
Tang, Z., Zhang, X., Wang, B., and Wei, F · 2024
Later among the works it cites.
Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving
Tong, Y., Zhang, X., Wang, R., Wu, R., and He, J · 2024
Later among the works it cites.
Chain-of-thought reasoning without prompting
Wang, X. and Zhou, D · 2024
Later among the works it cites.
Wen, K., Zhang, H., Lin, H., and Zhang, J · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Distilling system 2 into system 1
Yu, P., Xu, J., Weston, J., and Kulikov, I · 2024
Later among the works it cites.
Towards a theoretical understanding of the’reversal curse’via training dynamics
Zhu, H., Huang, B., Zhang, S., Jordan, M., Jiao, J., Tian, Y., and Russell, S · 2024
Later among the works it cites.
Galore 2: Large-scale llm pre-training by gradient low-rank projection
Su, D., Gu, A., Xu, J., Tian, Y., and Zhao, J · 2025
Closest in time.