Fetching the paper…
Reading the bibliography…
This paper introduces the MCT Self-Refine (MCTSr) algorithm, an innovative integration of Large Language Models (LLMs) with Monte Carlo Tree Search (MCTS), designed to enhance performance in complex mathematical reasoning tasks.
Monte-carlo tree search: A new framework for game ai
Chaslot, G., Bakkes, S. C. J., Szita, I., and Spronck, P. (2008) · 2008
Earlier work this paper cites.
Information-theoretic regret bounds for gaussian process optimization in the bandit setting
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. W. (2009) · 2009
Earlier work this paper cites.
A survey of monte carlo tree search methods
Browne, C., Powley, E. J., Whitehouse, D., Lucas, S. M. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Liebana, D. P., Samothrakis, S., and Colton, S. (2012) · 2012
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. (2021) · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Pal: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G. (2022) · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., hsin Chi, E. H., Xia, F., Le, Q., and Zhou, D. (2022) · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023) · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. (2023) · 2023
Earlier work this paper cites.
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., and Liu, T. (2023) · 2023
Earlier work this paper cites.
General method for solving four types of sat problems
Li, A., Han, C., Guo, T., Li, H., and Li, B. (2023) · 2023
Earlier work this paper cites.
Drugchat: Towards enabling chatgpt-like capabilities on drug molecule graphs
Liang, Y., Zhang, R., Zhang, L., and Xie, P. (2023) · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. (2023) · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D. (2023) · 2023
Cited alongside, same era.
Llm-rec: Personalized recommendation via prompting large language models
Lyu, H., Jiang, S., Zeng, H., Xia, Y., and Luo, J. (2023) · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z., Yuan, H., Li, C., Dong, G., Tan, C., and Zhou, C. (2023) · 2023
Later among the works it cites.
AGI odyssey
AGI Odyssey (2024) · 2024
Closest in time.
AIME problem set: 1983-2024
AIME (2024) · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, A. (2024) · 2024
Closest in time.
Alphamath almost zero: process supervision without process
Chen, G., Liao, M., Li, C., and Fan, K. (2024) · 2024
Closest in time.
Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Welleck, S., Majumder, B. P., Gupta, S., Yazdanbakhsh, A., and Clark, P. (2023) · 2023
Cited alongside, same era.
Monte-carlo tree search for multi-agent pathfinding: Preliminary results
Pitanov, Y., Skrynnik, A., Andreychuk, A., Yakovlev, K., and Panov, A. (2023) · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Cited alongside, same era.
Xu, H. (2023) · 2023
Cited alongside, same era.
Yang, F. (2023) · 2023
Cited alongside, same era.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W. (2023) · 2023
Cited alongside, same era.
Closest in time.
Introducing meta llama 3: The most capable openly available LLM to date
Meta AI (2024) · 2024
Closest in time.
Papers with code - GSM8k benchmark (arithmetic reasoning)
Papers with Code (2024) · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al. (2024) · 2024
Closest in time.
Vagadia, H., Chopra, M., Barnawal, A., Banerjee, T., Tuli, S., Chakraborty, S., and Paul, R. (2024) · 2024
Closest in time.