Fetching the paper…
Reading the bibliography…
Mathematical reasoning presents significant challenges for large language models (LLMs).
Nat. 529 , 484–489. URL: https://doi.org/10.1038/nature16961 . doi: 10.1038/NATURE16961
Silver, D., Huang, A., Maddison, C.J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T.P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016). Mastering the game of go with deep neural networks and tree search · 2016
Earlier work this paper cites.
Nat. 550 , 354–359. URL: https://doi.org/10.1038/nature24270 . doi: 10.1038/NATURE24270
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T.P., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. (2017). Mastering the game of go without human knowledge · 2017
Earlier work this paper cites.
IEEE Trans. Big Data 7 , 535–547. URL: https://doi.org/10.1109/TBDATA.2019.2921572 . doi: 10.1109/TBDATA.2019.2921572
Johnson, J., Douze, M., and Jégou, H. (2021). Billion-scale similarity search with gpus · 2019
Earlier work this paper cites.
URL: https://dl.acm.org/doi/abs/10.5555/3495724.3496517
Lewis, P.S.H., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks · 2020
Earlier work this paper cites.
CoRR abs/2110.14168 . URL: https://arxiv.org/abs/2110.14168 . arXiv:2110.14168
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. (2021). Training verifiers to solve math word problems · 2021
Earlier work this paper cites.
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021). Measuring mathematical problem solving with the MATH dataset · 2021
Earlier work this paper cites.
URL: https://dl.acm.org/doi/10.5555/3600270.3601883
Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., and Iwasawa, Y. (2022). Large language models are zero-shot reasoners · 2022
Earlier work this paper cites.
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. (2022). Self-critiquing models for assisting human evaluators · 2022
Earlier work this paper cites.
arXiv preprint arXiv:2211.12588
Chen, W., Ma, X., Wang, X., and Cohen, W.W. (2022). Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks · 2022
Earlier work this paper cites.
In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net
Fu, Y., Peng, H., Sabharwal, A., Clark, P., and Khot, T. (2023). Complexity-based prompting for multi-step reasoning · 2023
Earlier work this paper cites.
In P.G. Bringas, H.P. García, de F.J.M. Pisón, F. Martínez-Álvarez, A.T. Lora, Á. Herrero, J.L. Calvo-Rolle, H. Quintián, and E. Corchado, eds. Hybrid Artificial Intelligent Systems - 18th International Conference, HAIS 2023, Salamanca, Spain, September 5-7, 2023, Proceedings vol. 14001 of Lecture Notes in Computer Science . Springer pp. 649–660
Pitanov, Y., Skrynnik, A., Andreychuk, A., Yakovlev, K.S., and Panov, A. (2023). Monte-carlo tree search for multi-agent pathfinding: Preliminary results · 2023
Earlier work this paper cites.
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., Millican, K., Silver, D., Petrov, S., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T.P., Lazaridou, A., Firat, O., Molloy, J., Isard, M., Barham, P.R., Hennigan, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Doherty, R., Collins, E., Meyer, C., Rutherford, E., Moreira, E., Ayoub, K., Goel, M., Tucker, G., Piqueras, E., Krikun, M., Barr, I., Savinov, N., Danihelka, I., Roelofs, B., White, A., Andreassen, A., von Glehn, T., Yagati, L., Kazemi, M., Gonzalez, L., Khalman, M., Sygnowski, J., and et al. (2023). Gemini: A family of highly capable multimodal models · 2023
Earlier work this paper cites.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P.S., Lachaux, M., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E.M., Subramanian, R., Tan, X.E., Tang, B., Taylor, R., Williams, A., Kuan, J.X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models · 2023
Earlier work this paper cites.
Artif. Intell. Rev. 56 , 2497–2562. URL: https://doi.org/10.1007/s10462-022-10228-y . doi: 10.1007/S10462-022-10228-Y
Swiechowski, M., Godlewski, K., Sawicki, B., and Mandziuk, J. (2023). Monte carlo tree search: a review of recent modifications and applications · 2023
Cited alongside, same era.
In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, eds. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2023a). Tree of thoughts: Deliberate problem solving with large language models · 2023
Cited alongside, same era.
pp. 8154–8173. URL: https://doi.org/10.18653/v1/2023.emnlp-main.507 . doi: 10.18653/V1/2023.EMNLP-MAIN.507
Hao, S., Gu, Y., Ma, H., Hong, J.J., Wang, Z., Wang, D.Z., and Hu, Z. (2023). Reasoning with language model is planning with world model · 2023
Cited alongside, same era.
Yang, F. (2023). An integrated framework integrating monte carlo tree search and supervised learning for train timetabling problem · 2023
Zhang, D., Huang, X., Zhou, D., Li, Y., and Ouyang, W. (2024). Accessing GPT-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b · 2024
Closest in time.
Aime problems 1983 to 2024. Kaggle
Zamil, P., and Rabby, G. (2024) · 2024
Closest in time.
Fang, M., Wan, X., Lu, F., Xing, F., and Zou, K. (2024). Mathodyssey: Benchmarking mathematical problem-solving skills in large language models using odyssey math data · 2024
Closest in time.
Vagadia, H., Chopra, M., Barnawal, A., Banerjee, T., Tuli, S., Chakraborty, S., and Paul, R. (2024). Phyplan: Compositional and adaptive physical task reasoning with physics-informed skill networks for robot manipulators · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Li, A., Han, C., Guo, T., Li, H., and Li, B. (2023). General method for solving four types of SAT problems · 2023
Cited alongside, same era.
Xu, H. (2023). No train still gain. unleash mathematical reasoning of large language models with monte carlo tree search guided by energy function · 2023
Cited alongside, same era.
URL: https://dl.acm.org/doi/10.5555/3666122.3667430
Dubois, Y., Li, C.X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T.B. (2023). Alpacafarm: A simulation framework for methods that learn from human feedback · 2023
Cited alongside, same era.
Mark. Sci. 43 , 709–722. URL: https://doi.org/10.1287/mksc.2023.0306 . doi: 10.1287/MKSC.2023.0306
Goli, A., and Singh, A. (2024). Frontiers: Can large language models capture human preferences? · 2023
Cited alongside, same era.
Mitchell, M., Palmarini, A.B., and Moskvichev, A. (2023). Comparing humans, gpt-4, and GPT-4V on abstraction and reasoning tasks · 2023
Cited alongside, same era.
Jiang, A.Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D.S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L.R., Lachaux, M., Stock, P., Scao, T.L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W.E. (2023). Mistral 7b · 2023
Cited alongside, same era.
URL: https://openreview.net/forum?id=wCFB37bzud4
Patel, A., Li, B., Rasooli, M.S., Constant, N., Raffel, C., and Callison-Burch, C. (2023). Bidirectional language models are also few-shot learners · 2023
Cited alongside, same era.
URL: https://openreview.net/forum?id=C4OpREezgj
Wan, Z., Feng, X., Wen, M., McAleer, S.M., Wen, Y., Zhang, W., and Wang, J. (2024). Alphazero-like tree-search can guide large language model decoding and training · 2024
Cited alongside, same era.
Chen, G., Liao, M., Li, C., and Fan, K. (2024). Alphamath almost zero: process supervision without process · 2024
Closest in time.
In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D. (2024). Evaluating large language models at evaluating instruction following · 2024
Closest in time.
Ong, K.T., Kwon, T., and Yeo, J. (2024). Large language models are self-taught reasoners: Enhancing LLM applications via tailored problem-solving demonstrations · 2024
Closest in time.
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P., Lomeli, M., Hosseini, L., and Jégou, H. (2024). The faiss library · 2024
Closest in time.
Salesforce AI Research Blog 3
Meng, R., Liu, Y., Joty, S.R., Xiong, C., Zhou, Y., and Yavuz, S. (2024). Sfrembedding-mistral: enhance text retrieval with transfer learning · 2024
Closest in time.
Learning to reason with LLMs.
OpenAI (2024) · 2024
Closest in time.
Zhong, T., Liu, Z., Pan, Y., Zhang, Y., Zhou, Y., Liang, S., Wu, Z., Lyu, Y., Shu, P., Yu, X., Cao, C., Jiang, H., Chen, H., Li, Y., Chen, J., Hu, H., Liu, Y., Zhao, H., Xu, S., Dai, H., Zhao, L., Zhang, R., Zhao, W., Yang, Z., Chen, J., Wang, P., Ruan, W., Wang, H., Zhao, H., Zhang, J., Ren, Y., Qin, S., Chen, T., Li, J., Zidan, A.H., Jahin, A., Chen, M., Xia, S., Holmes, J., Zhuang, Y., Wang, J., Xu, B., Xia, W., Yu, J., Tang, K., Yang, Y., Sun, B., Yang, T., Lu, G., Wang, X., Chai, L., Li, H., Lu, J., Sun, L., Zhang, X., Ge, B., Hu, X., Zhang, L., Zhou, H., Zhang, L., Zhang, S., Liu, N., Jiang, B., Kong, L., Xiang, Z., Ren, Y., Liu, J., Jiang, X., Bao, Y., Zhang, W., Li, X., Li, G., Liu, W., Shen, D., Sikora, A., Zhai, X., Zhu, D., and Liu, T. (2024). Evaluation of openai o1: Opportunities and challenges of AGI · 2024
Closest in time.
arXiv preprint arXiv:2501.12948
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X. et al. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning · 2025
Closest in time.