Fetching the paper…
Reading the bibliography…
Self-correction is a novel method that can stimulate the potential reasoning abilities of large language models (LLMs).
Analysing Mathematical Reasoning Abilities of Neural Models
Saxton, D.; Grefenstette, E.; Hill, F.; and Kohli, P. 2019 · 2019
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021 · 2021
Earlier work this paper cites.
Measuring Mathematical Problem Solving With the MATH Dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Are NLP Models Really Able to Solve Simple Math Word Problems?
Patel, A.; Bhattamishra, S.; and Goyal, N. 2021 · 2021
Earlier work this paper cites.
Evidence-Based Factual Error Correction
Thorne, J.; and Vlachos, A. 2021 · 2021
Earlier work this paper cites.
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Suzgun, M.; Scales, N.; Schärli, N.; Gehrmann, S.; Tay, Y.; Chung, H. W.; Chowdhery, A.; Le, Q. V.; Chi, E. H.; Zhou, D.; and Wei, J. 2022 · 2022
Earlier work this paper cites.
LM vs LM: Detecting Factual Errors via Cross Examination
Cohen, R.; Hamri, M.; Geva, M.; and Globerson, A. 2023 · 2023
Earlier work this paper cites.
Improving Factuality and Reasoning in Language Models through Multiagent Debate
Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J. B.; and Mordatch, I. 2023 · 2023
Earlier work this paper cites.
Active Retrieval Augmented Generation
Jiang, Z.; Xu, F.; Gao, L.; Sun, Z.; Liu, Q.; Dwivedi-Yu, J.; Yang, Y.; Callan, J.; and Neubig, G. 2023 · 2023
Earlier work this paper cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C. H.; Gonzalez, J. E.; Zhang, H.; and Stoica, I. 2023 · 2023
Earlier work this paper cites.
Let’s Verify Step by Step
Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023 · 2023
Earlier work this paper cites.
Self-Refine: Iterative Refinement with Self-Feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; Gupta, S.; Majumder, B. P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P. 2023 · 2023
Earlier work this paper cites.
Mistral 7B
Mistral AI team. 2023 · 2023
Cited alongside, same era.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023 · 2023
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2023 · 2023
Cited alongside, same era.
Step-Level Value Preference Optimization for Mathematical Reasoning
Chen, G.; Liao, M.; Li, C.; and Fan, K. 2024 · 2024
Cited alongside, same era.
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
Kamoi, R.; Zhang, Y.; Zhang, N.; Han, J.; and Zhang, R. 2024 · 2024
Cited alongside, same era.
PRD: Peer Rank and Discussion Improve Large Language Model Based Evaluations
Introducing Qwen2-Math | Qwen
Qwen Team. 2024 · 2024
Closest in time.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Zhang, M.; Li, Y. K.; Wu, Y.; and Guo, D. 2024 · 2024
Closest in time.
A Survey of Reasoning with Foundation Models
Sun, J.; Zheng, C.; Xie, E.; Liu, Z.; Chu, R.; Qiu, J.; Xu, J.; Ding, M.; Li, H.; Geng, M.; Wu, Y.; Wang, W.; Chen, J.; Yin, Z.; Ren, X.; Fu, J.; He, J.; Yuan, W.; Liu, Q.; Liu, X.; Li, Y.; Dong, H.; Cheng, Y.; Zhang, M.; Heng, P. A.; Dai, J.; Luo, P.; Wang, J.; Wen, J.-R.; Qiu, X.; Guo, Y.; Xiong, H.; Liu, Q.; and Li, Z. 2024 · 2024
Closest in time.
LLMs Cannot Find Reasoning Errors, but Can Correct Them given the Error Location
Tyen, G.; Mansoor, H.; Carbune, V.; Chen, P.; and Mak, T. 2024 · 2024
Closest in time.
Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations
Wang, P.; Li, L.; Shao, Z.; Xu, R. X.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, R.; Patel, T.; and Du, X. 2024 · 2024
Cited alongside, same era.
Large Language Models Have Intrinsic Self-Correction Ability
Liu, D.; Nassereldine, A.; Yang, Z.; Xu, C.; Hu, Y.; Li, J.; Kumar, U.; Lee, C.; and Xiong, J. 2024 · 2024
Cited alongside, same era.
Augmenting math word problems via iterative question composing
Liu, H.; and Yao, A. C.-C. 2024 · 2024
Cited alongside, same era.
Introducing Meta Llama 3: The Most Capable Openly Available LLM to Date
Meta AI. 2024 · 2024
Cited alongside, same era.
Automatically Correcting Large Language Models: Surveying the Landscape of Diverse Automated Correction Strategies
Pan, L.; Saxon, M.; Xu, W.; Nathani, D.; Wang, X.; and Wang, W. Y. 2024 · 2024
Cited alongside, same era.
REFINER: Reasoning Feedback on Intermediate Representations
Paul, D.; Ismayilzada, M.; Peyrard, M.; Borges, B.; Bosselut, A.; West, R.; and Faltings, B. 2024 · 2024
Cited alongside, same era.
Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models
Puerto, H.; Chubakov, T.; Zhu, X.; Madabushi, H. T.; and Gurevych, I. 2024 · 2024
Cited alongside, same era.
Chain-of-Thought Reasoning Without Prompting
Wang, X.; and Zhou, D. 2024 · 2024
Closest in time.
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
Ying, H.; Zhang, S.; Li, L.; Zhou, Z.; Shao, Y.; Fei, Z.; Ma, Y.; Hong, J.; Liu, K.; Wang, Z.; Wang, Y.; Wu, Z.; Li, S.; Zhou, F.; Liu, H.; Zhang, S.; Zhang, W.; Yan, H.; Qiu, X.; Wang, J.; Chen, K.; and Lin, D. 2024 · 2024
Closest in time.
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Yu, L.; Jiang, W.; Shi, H.; Yu, J.; Liu, Z.; Zhang, Y.; Kwok, J. T.; Li, Z.; Weller, A.; and Liu, W. 2024 · 2024
Closest in time.
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
Zheng, T.; Zhang, G.; Shen, T.; Liu, X.; Lin, B. Y.; Fu, J.; Chen, W.; and Yue, X. 2024 · 2024
Closest in time.
JiuZhang3. 0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models
Zhou, K.; Zhang, B.; Wang, J.; Chen, Z.; Zhao, W. X.; Sha, J.; Sheng, Z.; Wang, S.; and Wen, J.-R. 2024 · 2024
Closest in time.
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
Zhu, X.; Li, J.; Liu, Y.; Ma, C.; and Wang, W. 2024 · 2024
Closest in time.