Fetching the paper…
Reading the bibliography…
We propose a framework for robust evaluation of reasoning capabilities of language models, using functional variants of benchmarks.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Quac : Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the AI2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? A new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2018
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2019
Earlier work this paper cites.
On the measure of intelligence, 2019
François Chollet · 2019
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
WINOGRANDE: an adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2019
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Ar-lsat: Investigating analytical reasoning of text, 2021
Wanjun Zhong, Siyuan Wang, Duyu Tang, Zenan Xu, Daya Guo, Jiahai Wang, Jian Yin, Ming Zhou, and Nan Duan · 2021
Earlier work this paper cites.
Talm: Tool augmented language models, 2022
Aaron Parisi, Yao Zhao, and Noah Fiedel · 2022
Earlier work this paper cites.
Learning by distilling context, 2022
Charlie Snell, Dan Klein, and Ruiqi Zhong · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them, 2022
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei · 2022
Earlier work this paper cites.
The clrs algorithmic reasoning benchmark, 2022
Petar Veličković, Adrià Puigdomènech Badia, David Budden, Razvan Pascanu, Andrea Banino, Misha Dashevskiy, Raia Hadsell, and Charles Blundell · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou · 2022
Earlier work this paper cites.
Teaching algorithmic reasoning via in-context learning, 2022
Hattie Zhou, Azade Nova, Hugo Larochelle, Aaron Courville, Behnam Neyshabur, and Hanie Sedghi · 2022
Earlier work this paper cites.
Long context prompting for claude 2.1, 2023
2023
Earlier work this paper cites.
Building the next generation of open-source and bilingual llms, 2023
01.AI · 2023
Earlier work this paper cites.
Rest meets react: Self-improvement for multi-step reasoning llm agent, 2023
Renat Aksitov, Sobhan Miryoosefi, Zonglin Li, Daliang Li, Sheila Babayan, Kavya Kopparapu, Zachary Fisher, Ruiqi Guo, Sushant Prakash, Pranesh Srinivasan, Manzil Zaheer, Felix Yu, and Sanjiv Kumar · 2023
Earlier work this paper cites.
Rest meets react: Self-improvement for multi-step reasoning llm agent, 2023
Renat Aksitov, Sobhan Miryoosefi, Zonglin Li, Daliang Li, Sheila Babayan, Kavya Kopparapu, Zachary Fisher, Ruiqi Guo, Sushant Prakash, Pranesh Srinivasan, Manzil Zaheer, Felix Yu, and Sanjiv Kumar · 2023
Earlier work this paper cites.
Introducing claude 2.1, 2023
Anthropic · 2023
Earlier work this paper cites.
Have llms advanced enough? a challenging problem solving benchmark for large language models, 2023
Daman Arora, Himanshu Gaurav Singh, and Mausam · 2023
Earlier work this paper cites.
Proofnet: Autoformalizing and formally proving undergraduate-level mathematics, 2023
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, and Jeremy Avigad · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Earlier work this paper cites.
Worldsense: A synthetic benchmark for grounded reasoning in large language models, 2023
Youssef Benchekroun, Megi Dervishi, Mark Ibrahim, Jean-Baptiste Gaya, Xavier Martinet, Grégoire Mialon, Thomas Scialom, Emmanuel Dupoux, Dieuwke Hupkes, and Pascal Vincent · 2023
Cited alongside, same era.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks, 2023
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen · 2023
Cited alongside, same era.
Teaching large language models to self-debug, 2023
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou · 2023
Cited alongside, same era.
The capacity for moral self-correction in large language models, 2023
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, Dawn Drain, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jackson Kernion, Jamie Kerr, Jared Mueller, Joshua Landau, Kamal Ndousse, Karina Nguyen, Liane Lovitt, Michael Sellitto, Nelson Elhage, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robert Lasenby, Robin Larson, Sam Ringer, Sandipan Kundu, Saurav Kadavath, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, Christopher Olah, Jack Clark, Samuel R. Bowman, and Jared Kaplan · 2023
Cited alongside, same era.
Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems, 2023
Kaya Stechly, Matthew Marquez, and Subbarao Kambhampati · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Gemini Team · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models, 2023
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions, 2023
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks, 2023
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pal: Program-aided language models, 2023
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Cited alongside, same era.
Roscoe: A suite of metrics for scoring step-by-step reasoning, 2023
Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz · 2023
Cited alongside, same era.
Large language models are reasoning teachers, 2023
Namgyu Ho, Laura Schmid, and Se-Young Yun · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet, 2023
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2023
Cited alongside, same era.
Mathprompter: Mathematical reasoning using large language models, 2023
Shima Imani, Liang Du, and Harsh Shrivastava · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?, 2023
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
xcodeeval: A large scale multilingual multitask benchmark for code understanding, generation, translation and retrieval, 2023
Mohammad Abdullah Matin Khan, M Saiful Bari, Xuan Long Do, Weishi Wang, Md Rizwan Parvez, and Shafiq Joty · 2023
Cited alongside, same era.
Task contamination: Language models may not be few-shot anymore, 2023
Changmao Li and Jeffrey Flanigan · 2023
Cited alongside, same era.
Pretraining data mixtures enable narrow model selection capabilities in transformer models, 2023
Steve Yadlowsky, Lyric Doshi, and Nilesh Tripuraneni · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Later among the works it cites.
React: Synergizing reasoning and acting in language models, 2023
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Later among the works it cites.
Self-taught optimizer (stop): Recursively self-improving code generation, 2023
Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai · 2023
Later among the works it cites.
R-tuning: Teaching large language models to refuse unknown questions, 2023
Hanning Zhang, Shizhe Diao, Yong Lin, Yi R. Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang · 2023
Later among the works it cites.
Planning with large language models for code generation, 2023
Shun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding, Joshua B. Tenenbaum, and Chuang Gan · 2023
Later among the works it cites.
Progressive-hint prompting improves reasoning in large language models, 2023
Chuanyang Zheng, Zhengying Liu, Enze Xie, Zhenguo Li, and Yu Li · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models, 2023
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi · 2023
Later among the works it cites.
Self-play fine-tuning converts weak language models to strong language models, 2024
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Closest in time.
Symbolicai: A framework for logic-based approaches combining generative models and solvers, 2024
Marius-Constantin Dinu, Claudiu Leoveanu-Condrei, Markus Holzleitner, Werner Zellinger, and Sepp Hochreiter · 2024
Closest in time.
Tora: A tool-integrated reasoning agent for mathematical problem solving, 2024
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen · 2024
Closest in time.
Cruxeval: A benchmark for code reasoning, understanding and execution, 2024
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I. Wang · 2024
Closest in time.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
The impact of reasoning step length on large language models, 2024
Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du · 2024
Closest in time.
Same task, more tokens: the impact of input length on the reasoning performance of large language models, 2024
Mosh Levy, Alon Jacoby, and Yoav Goldberg · 2024
Closest in time.
Divide and conquer for large language models reasoning, 2024
Zijie Meng, Yan Zhang, Zhaopeng Feng, Yang Feng, Gaoang Wang, Joey Tianyi Zhou, Jian Wu, and Zuozhu Liu · 2024
Closest in time.
A survey of reasoning with foundation models, 2024
Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Liu, Ruihang Chu, Jianing Qiu, Jiaqi Xu, Mingyu Ding, Hongyang Li, Mengzhe Geng, Yue Wu, Wenhai Wang, Junsong Chen, Zhangyue Yin, Xiaozhe Ren, Jie Fu, Junxian He, Wu Yuan, Qi Liu, Xihui Liu, Yu Li, Hao Dong, Yu Cheng, Ming Zhang, Pheng Ann Heng, Jifeng Dai, Ping Luo, Jingdong Wang, Ji-Rong Wen, Xipeng Qiu, Yike Guo, Hui Xiong, Qun Liu, and Zhenguo Li · 2024
Closest in time.
Llms cannot find reasoning errors, but can correct them!, 2024
Gladys Tyen, Hassan Mansoor, Victor Cărbune, Peter Chen, and Tony Mak · 2024
Closest in time.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations, 2024
Peiyi Wang, Lei Li, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen, Y. Wu, and Zhifang Sui · 2024
Closest in time.
Understanding the reasoning ability of language models from the perspective of reasoning paths aggregation, 2024
Xinyi Wang, Alfonso Amayuelas, Kexun Zhang, Liangming Pan, Wenhu Chen, and William Yang Wang · 2024
Closest in time.
Chain-of-thought reasoning without prompting, 2024
Xuezhi Wang and Denny Zhou · 2024
Closest in time.
If llm is the wizard, then code is the wand: A survey on how code empowers large language models to serve as intelligent agents, 2024
Ke Yang, Jiateng Liu, John Wu, Chaoqi Yang, Yi R. Fung, Sha Li, Zixuan Huang, Xu Cao, Xingyao Wang, Yiquan Wang, Heng Ji, and Chengxiang Zhai · 2024
Closest in time.
Do large language models latently perform multi-hop reasoning?, 2024
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel · 2024
Closest in time.
Benchmarking llms via uncertainty quantification, 2024
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
Large language models as an indirect reasoner: Contrapositive and contradiction for automated reasoning, 2024
Yanfang Zhang, Yiliu Sun, Yibing Zhan, Dapeng Tao, Dacheng Tao, and Chen Gong · 2024
Closest in time.
Self-discover: Large language models self-compose reasoning structures, 2024
Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, Denny Zhou, Swaroop Mishra, and Huaixiu Steven Zheng · 2024
Closest in time.