Fetching the paper…
Reading the bibliography…
Recent advancements in Large Language Models (LLMs) have showcased striking results on existing logical reasoning benchmarks, with some models even surpassing human performance.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Xiaodong Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Self-supervised bug detection and repair
Miltiadis Allamanis, Henry Jackson-Flux, and Marc Brockschmidt · 2021
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Explaining the efficacy of counterfactually augmented data
Divyansh Kaushik, Amrith Setlur, Eduard Hovy, and Zachary C Lipton · 2021
Earlier work this paper cites.
Are nlp models really able to solve simple math word problems?
Arkil Patel, S. Bhattamishra, and Navin Goyal · 2021
Earlier work this paper cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu · 2021
Earlier work this paper cites.
Causalqa: A benchmark for causal question answering
Alexander Bondarenko, Magdalena Wolska, Stefan Heindorf, Lukas Blübaum, Axel-Cyrille Ngonga Ngomo, Benno Stein, Pavel Braslavski, Matthias Hagen, and Martin Potthast · 2022
Earlier work this paper cites.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen · 2022
Earlier work this paper cites.
Language models show human-like content effects on reasoning
Ishita Dasgupta, Andrew Kyle Lampinen, Stephanie C. Y. Chan, Antonia Creswell, Dharshan Kumaran, James L. McClelland, and Felix Hill · 2022
Earlier work this paper cites.
Pal: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2022
Earlier work this paper cites.
Robustlr: A diagnostic benchmark for evaluating logical robustness of deductive reasoners
Soumya Sanyal, Zeyi Liao, and Xiang Ren · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Huai hsin Chi, and Denny Zhou · 2022
Cited alongside, same era.
The reversal curse: Llms trained on "a is b" fail to learn "b is a"
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2023
Cited alongside, same era.
Aniruddha Deb, Neeva Oza, Sarthak Singla, Dinesh Khandelwal, Dinesh Garg, and Parag Singla · 2023
Cited alongside, same era.
Clomo: Counterfactual logical modification with large language models
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Huai hsin Chi, Nathanael Scharli, and Denny Zhou · 2023
Later among the works it cites.
Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks, 2023
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Later among the works it cites.
Mr-gsm8k: A meta-reasoning benchmark for large language model evaluation
Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiang, and Jiaya Jia · 2023
Later among the works it cites.
Mathattack: Attacking large language models towards math solving ability
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinya Huang, Ruixin Hong, Hongming Zhang, Wei Shao, Zhicheng YANG, Dong Yu, Changshui Zhang, Xiaodan Liang, and Linqi Song · 2023
Cited alongside, same era.
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Zekun Li, Baolin Peng, Pengcheng He, and Xifeng Yan · 2023
Cited alongside, same era.
A symbolic framework for evaluating mathematical reasoning and generalisation with transformers
Jordan Meadows, Marco Valentino, Damien Teney, and André Freitas · 2023
Cited alongside, same era.
Sync: A structurally guided hard negative curricula for generalizable neural code search
Atharva Naik, Soumitra Das, Jyothi Vedurada, and Somak Aditya · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, I. Evtimov, Joanna Bitton, Manish P Bhatt, Cristian Cantón Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre D’efossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2023
Cited alongside, same era.
Arb: Advanced reasoning benchmark for large language models
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla, Pranav Tadepalli, Paula Vidas, Alexander Kranias, John J. Nay, Kshitij Gupta, and Aran Komatsuzaki · 2023
Cited alongside, same era.
Zihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao, Jianan Ye, Wei Liu, Wei Wang, Xiaowei Huang, and Kaizhu Huang · 2023
Later among the works it cites.
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin · 2024
Closest in time.
Beyond accuracy: Evaluating the reasoning behavior of large language models - a survey
Philipp Mondorf and Barbara Plank · 2024
Closest in time.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
Saurabh Srivastava, B AnnaroseM, V AntoP, Shashank Menon, Ajay Sukumar, T AdwaithSamod, Alan Philipose, Stevin Prince, and Sooraj Thomas · 2024
Closest in time.
A & b == b & a: Triggering logical reasoning failures in large language models
Yuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan, Jen tse Huang, Pinjia He, Wenxiang Jiao, and Michael R. Lyu · 2024
Closest in time.
Benchmark self-evolving: A multi-agent framework for dynamic llm evaluation
Siyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu Wei, and Xuanjing Huang · 2024
Closest in time.
Evaluating mathematical reasoning beyond accuracy
Shijie Xia, Xuefeng Li, Yixin Liu, Tongshuang Wu, and Pengfei Liu · 2024
Closest in time.