Fetching the paper…
Reading the bibliography…
Benchmarks are critical for measuring Large Language Model (LLM) reasoning capabilities.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
PAL: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 · 2023
Earlier work this paper cites.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R Narasimhan. 2023 · 2023
Earlier work this paper cites.
GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models
Anonymous. 2024 · 2024
Earlier work this paper cites.
Introducing claude 3.5 sonnet
Anthropic. 2024 · 2024
Earlier work this paper cites.
Premise order matters in reasoning with large language models
Xinyun Chen, Ryan A. Chi, Xuezhi Wang, and Denny Zhou. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, , and Others. 2024 · 2024
Cited alongside, same era.
Alex Havrilla, Andrew Dai, Laura O’Mahony, Koen Oostermeijer, Vera Zisler, Alon Albalak, Fabrizio Milo, Sharath Chandra Raparthy, Kanishk Gandhi, Baber Abbasi, Duy Phung, Maia Iyer, Dakota Mahan, Chase Blagden, Srishti Gureja, Mohammed Hamdy, Wen-Ding Li, Giovanni Paolini, Pawan Sasanka Ammanamanchi, and Elliot Meyerson. 2024 · 2024
Cited alongside, same era.
Not all LLM reasoners are created equal
Arian Hosseini, Alessandro Sordoni, Daniel Kenji Toyama, Aaron Courville, and Rishabh Agarwal. 2024 · 2024
Cited alongside, same era.
Hello gpt-4o
OpenAI. 2024a · 2024
Closest in time.
Learning to reason with llms (openai o1)
OpenAI. 2024b · 2024
Closest in time.
OpenAI, Josh Achiam, and Others. 2024 · 2024
Closest in time.
Arithmetic reasoning on gsm8k
paperswithcode. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, , and Others. 2024 · 2024
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Li, B. Jasani, P. Tang, and S. Ghadar. 2024 · 2024
Cited alongside, same era.
Best practices and lessons learned on synthetic data
Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, and Andrew M. Dai. 2024 · 2024
Cited alongside, same era.
On leakage of code generation evaluation datasets
Alexandre Matton, Tom Sherborne, Dennis Aumiller, Elena Tommasone, Milad Alizadeh, Jingyi He, Raymond Ma, Maxime Voisin, Ellen Gilsenan-McMahon, and Matthias Gallé. 2024 · 2024
Cited alongside, same era.
A careful examination of large language model performance on grade school arithmetic
Hugh Zhang, Jeff Da, Dean Lee, Vaughn Robinson, Catherine Wu, Will Song, Tiffany Zhao, Pranav Raja, Dylan Slack, Qin Lyu, Sean Hendryx, Russell Kaplan, Michele Lunati, and Summer Yue. 2024a
Cited in the paper.
Training language models with syntactic data generation
Yifan Zhang, Yifan Luo, Yang Yuan, and Andrew Chi-Chih Yao. 2024b
Cited in the paper.
Physics of language models: Part 2.1, grade-school math and the hidden reasoning process
Tian Ye, Zicheng Xu, Yuanzhi Li, and Zeyuan Allen-Zhu. 2024 · 2024
Closest in time.