Fetching the paper…
Reading the bibliography…
With current state-of-the-art approaches aimed at enhancing the reasoning capabilities of Large Language Models(LLMs) through iterative preference learning inspired by AlphaZero, we propose to further enhance the step-wise reasoning capabilities through intrinsic self-correction to some extent.
Efficient selectivity and backup operators in Monte-Carlo tree search
Coulom, R. 2006 · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L.; and Szepesvári, C. 2006 · 2006
Earlier work this paper cites.
Multi-armed bandits with episode context
Rosin, C. D. 2011 · 2011
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D.; Hubert, T.; Schrittwieser, J.; Antonoglou, I.; Lai, M.; Guez, A.; Lanctot, M.; Sifre, L.; Kumaran, D.; Graepel, T.; et al. 2017 · 2017
Earlier work this paper cites.
Monte-Carlo tree search as regularized policy optimization
Grill, J.-B.; Altché, F.; Tang, Y.; Hubert, T.; Valko, M.; Antonoglou, I.; and Munos, R. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Offline rl for natural language generation with implicit language q learning
Snell, C.; Kostrikov, I.; Su, Y.; Yang, M.; and Levine, S. 2022 · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J.; Kushman, N.; Kumar, R.; Song, F.; Siegel, N.; Wang, L.; Creswell, A.; Irving, G.; and Higgins, I. 2022 · 2022
Earlier work this paper cites.
Generating sequences by learning to self-correct
Welleck, S.; Lu, X.; West, P.; Brahman, F.; Shen, T.; Khashabi, D.; and Choi, Y. 2022 · 2022
Earlier work this paper cites.
Solving math word problems via cooperative reasoning induced language models
Zhu, X.; Wang, J.; Zhang, L.; Zhang, Y.; Gan, R.; Zhang, J.; and Yang, Y. 2022 · 2022
Earlier work this paper cites.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
Akyürek, A. F.; Akyürek, E.; Madaan, A.; Kalyan, A.; Clark, P.; Wijaya, D.; and Tandon, N. 2023 · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Hao, S.; Gu, Y.; Ma, H.; Hong, J. J.; Wang, Z.; Wang, D. Z.; and Hu, Z. 2023 · 2023
Earlier work this paper cites.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Earlier work this paper cites.
Making language models better reasoners with step-aware verifier
Li, Y.; Lin, Z.; Zhang, S.; Fu, Q.; Chen, B.; Lou, J.-G.; and Chen, W. 2023 · 2023
Cited alongside, same era.
Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023 · 2023
Cited alongside, same era.
Making ppo even better: Value-guided monte-carlo tree search decoding
Liu, J.; Cohen, A.; Pasunuru, R.; Choi, Y.; Hajishirzi, H.; and Celikyilmaz, A. 2023 · 2023
Cited alongside, same era.
Refiner: Reasoning feedback on intermediate representations
Paul, D.; Ismayilzada, M.; Peyrard, M.; Borges, B.; Bosselut, A.; West, R.; and Faltings, B. 2023 · 2023
Cited alongside, same era.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Training language models to self-correct via reinforcement learning
Kumar, A.; Zhuang, V.; Agarwal, R.; Su, Y.; Co-Reyes, J. D.; Singh, A.; Baumli, K.; Iqbal, S.; Bishop, C.; Roelofs, R.; et al. 2024 · 2024
Closest in time.
The Llama 3 Herd of Models
Llama Team, A. . M. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2024 · 2024
Closest in time.
Recursive introspection: Teaching language model agents how to self-improve
Qu, Y.; Zhang, T.; Garg, N.; and Kumar, A. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shinn, N.; Labash, B.; and Gopinath, A. 2023 · 2023
Cited alongside, same era.
Fine-grained human feedback gives better rewards for language model training
Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N. A.; Ostendorf, M.; and Hajishirzi, H. 2023 · 2023
Cited alongside, same era.
Decomposition enhances reasoning via self-evaluation guided decoding
Xie, Y.; Kawaguchi, K.; Zhao, Y.; Zhao, X.; Kan, M.-Y.; He, J.; and Xie, Q. 2023 · 2023
Cited alongside, same era.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L.; Jiang, W.; Shi, H.; Yu, J.; Liu, Z.; Zhang, Y.; Kwok, J. T.; Li, Z.; Weller, A.; and Liu, W. 2023 · 2023
Cited alongside, same era.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Ahmadian, A.; Cremer, C.; Gallé, M.; Fadaee, M.; Kreutzer, J.; Pietquin, O.; Üstün, A.; and Hooker, S. 2024 · 2024
Cited alongside, same era.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Cited alongside, same era.
Stop regressing: Training value functions via classification for scalable deep rl
Farebrother, J.; Orbay, J.; Vuong, Q.; Taïga, A. A.; Chebotar, Y.; Xiao, T.; Irpan, A.; Levine, S.; Castro, P. S.; Faust, A.; et al. 2024 · 2024
Cited alongside, same era.
Glore: When, where, and how to improve llm reasoning via global and local refinements
Havrilla, A.; Raparthy, S.; Nalmpantis, C.; Dwivedi-Yu, J.; Zhuravinskyi, M.; Hambro, E.; and Raileanu, R. 2024 · 2024
Cited alongside, same era.
Shani, L.; Rosenberg, A.; Cassel, A.; Lang, O.; Calandriello, D.; Zipori, A.; Noga, H.; Keller, O.; Piot, B.; Szpektor, I.; et al. 2024 · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y.; Wu, Y.; et al. 2024 · 2024
Closest in time.
Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving
Tong, Y.; Zhang, X.; Wang, R.; Wu, R.; and He, J. 2024 · 2024
Closest in time.
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Xie, Y.; Goyal, A.; Zheng, W.; Kan, M.-Y.; Lillicrap, T. P.; Kawaguchi, K.; and Shieh, M. 2024 · 2024
Closest in time.
Building math agents with multi-turn iterative preference learning
Xiong, W.; Shi, C.; Shen, J.; Rosenberg, A.; Qin, Z.; Calandriello, D.; Khalman, M.; Joshi, R.; Piot, B.; Saleh, M.; et al. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2024 · 2024
Closest in time.
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
Zhang, Y.; Khalifa, M.; Logeswaran, L.; Kim, J.; Lee, M.; Lee, H.; and Wang, L. 2024 · 2024
Closest in time.
Archer: Training language model agents via hierarchical multi-turn rl
Zhou, Y.; Zanette, A.; Pan, J.; Levine, S.; and Kumar, A. 2024 · 2024
Closest in time.