Fetching the paper…
Reading the bibliography…
In recent years, there has been remarkable progress in leveraging Language Models (LMs), encompassing Pre-trained Language Models (PLMs) and Large-scale Language Models (LLMs), within the domain of mathematics.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
From ’F’ to ’a’ on the N.Y. Regents Science Exams: An Overview of the Aristo Project
Peter Clark, Oren Etzioni, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, Niket Tandon, Sumithra Bhakthavatsalam, Dirk Groeneveld, Michal Guerquin, and Michael Schmitz. 2021 · 1909
Earlier work this paper cites.
On Faithfulness and Factuality in Abstractive Summarization. In ACL . 1906–1919
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Empirical Explorations of the Logic Theory Machine: A Case Study in Heuristic. 218–230
A. Newell, J. C. Shaw, and H. A. Simon. 1957 · 1957
Earlier work this paper cites.
Empirical explorations of the geometry theorem machine. In western joint IRE-AIEE-ACM computer conference . 143–149
Herbert Gelernter, James R Hansen, and Donald W Loveland. 1960 · 1960
Earlier work this paper cites.
Computers and thought . Vol. 7
Edward A Feigenbaum, Julian Feldman, et al · 1963
Earlier work this paper cites.
Natural language input for a computer problem solving system
Daniel Bobrow et al · 1964
Earlier work this paper cites.
The socratic method
Leonard Nelson. 1980 · 1980
Earlier work this paper cites.
An integrated model of skill in solving elementary word problems
Diane J Briars and Jill H Larkin. 1984 · 1984
Earlier work this paper cites.
Understanding and solving arithmetic word problems: A computer simulation
Charles R Fletcher. 1985 · 1985
Earlier work this paper cites.
An overview of the Mizar project. In Proceedings of the 1992 Workshop on Types for Proofs and Programs . 311–330
Piotr Rudnicki. 1992 · 1992
Earlier work this paper cites.
A dialogue about Socratic teaching
Peggy Cooper Davis and Elizabeth Ehrenfest Steinglass. 1997 · 1997
Earlier work this paper cites.
Monte-carlo tree search: A new framework for game ai. In AAAI , Vol. 4. 216–217
Guillaume Chaslot, Sander Bakkes, Istvan Szita, and Pieter Spronck. 2008 · 2008
Earlier work this paper cites.
A brief overview of HOL4. In TPHOLs . 28–32
Konrad Slind and Michael Norrish. 2008 · 2008
Earlier work this paper cites.
The isabelle framework. In TPHOLs . 33–38
Makarius Wenzel, Lawrence C Paulson, and Tobias Nipkow. 2008 · 2008
Earlier work this paper cites.
Socratic teaching and Socratic method
Thomas C Brickhouse and Nicholas D Smith. 2009 · 2009
Earlier work this paper cites.
Generative Language Modeling for Automated Theorem Proving
Stanislas Polu and Ilya Sutskever. 2020a · 2009
Earlier work this paper cites.
Interactive theorem proving and program development: Coq’Art: the calculus of inductive constructions
Yves Bertot and Pierre Castéran. 2013 · 2013
Earlier work this paper cites.
A machine-checked proof of the odd order theorem. In ITP . 163–179
Georges Gonthier, Andrea Asperti, Jeremy Avigad, Yves Bertot, Cyril Cohen, François Garillot, Stéphane Le Roux, Assia Mahboubi, Russell O’Connor, Sidi Ould Biha, et al · 2013
Earlier work this paper cites.
Learning to Solve Arithmetic Word Problems with Verb Categorization. In EMNLP . 523–533
Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014 · 2014
Earlier work this paper cites.
Learning to Automatically Solve Algebra Word Problems. In ACL . 271–281
Nate Kushman, Yoav Artzi, Luke Zettlemoyer, and Regina Barzilay. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In ICCV . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Four Decades of Mizar: Foreword
Adam Grabowski, Artur Korniłowicz, and Adam Naumowicz. 2015 · 2015
Earlier work this paper cites.
Parsing algebraic word problems into equations
Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015 · 2015
Earlier work this paper cites.
Solving General Arithmetic Word Problems. In EMNLP . 1743–1752
Subhro Roy and Dan Roth. 2015 · 2015
Earlier work this paper cites.
Reasoning about quantities in natural language
Subhro Roy, Tim Vieira, and Dan Roth. 2015 · 2015
Earlier work this paper cites.
Solving Geometry Problems: Combining Text and Diagram Interpretation. In EMNLP . 1466–1476
Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, and Clint Malcolm. 2015 · 2015
Earlier work this paper cites.
Automatically solving number word problems by semantic parsing and reasoning. In EMNLP . 1132–1142
Shuming Shi, Yuehui Wang, Chin-Yew Lin, Xiaojiang Liu, and Yong Rui. 2015 · 2015
Earlier work this paper cites.
Draw: A challenging and diverse algebra word problem set
Shyam Upadhyay and Ming-Wei Chang. 2015 · 2015
Earlier work this paper cites.
Learn to solve algebra word problems using quadratic programming. In EMNLP . 817–822
Lipu Zhou, Shuaixiang Dai, and Liwei Chen. 2015 · 2015
Earlier work this paper cites.
How well do computers solve math word problems? large-scale dataset construction and evaluation. In ACL . 887–896
Danqing Huang, Shuming Shi, Chin-Yew Lin, Jian Yin, and Wei-Ying Ma. 2016 · 2016
Earlier work this paper cites.
Deepmath-deep sequence models for premise selection
Geoffrey Irving, Christian Szegedy, Alexander A Alemi, Niklas Eén, François Chollet, and Josef Urban. 2016 · 2016
Earlier work this paper cites.
MAWPS: A math word problem repository. In NAACL . 1152–1157
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016 · 2016
Earlier work this paper cites.
Learning to use formulas to solve simple arithmetic problems. In ACL . 2144–2153
Arindam Mitra and Chitta Baral. 2016 · 2016
Earlier work this paper cites.
Numerically Grounded Language Models for Semantic Error Correction. In EMNLP . 987–992
Georgios Spithourakis, Isabelle Augenstein, and Sebastian Riedel. 2016 · 2016
Earlier work this paper cites.
Verb Physics: Relative Physical Knowledge of Actions and Objects. In ACL . 266–276
Maxwell Forbes and Yejin Choi. 2017 · 2017
Earlier work this paper cites.
A formal proof of the Kepler conjecture. In Forum of mathematics, Pi , Vol. 5. e2
Thomas Hales, Mark Adams, Gertrud Bauer, Tat Dat Dang, John Harrison, Hoang Le Truong, Cezary Kaliszyk, Victor Magron, Sean McLaughlin, Tat Thang Nguyen, et al · 2017
Earlier work this paper cites.
HolStep: A Machine Learning Dataset for Higher-order Logic Theorem Proving. In ICLR
Cezary Kaliszyk, François Chollet, and Christian Szegedy. 2017 · 2017
Earlier work this paper cites.
Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems. In ACL . 158–167
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
Unit dependency graph and its application to arithmetic word problem solving. In AAAI , Vol. 31
Subhro Roy and Dan Roth. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Annotating Derivations: A New Evaluation Strategy and Dataset for Algebra Word Problems. In EACL . 494–504
Shyam Upadhyay and Ming-Wei Chang. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Deep neural solver for math word problems. In EMNLP . 845–854
Yan Wang, Xiaojiang Liu, and Shuming Shi. 2017 · 2017
Earlier work this paper cites.
Improving Language Understanding by Generative Pre-Training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
Mapping to declarative knowledge for word problem solving
Subhro Roy and Dan Roth. 2018 · 2018
Earlier work this paper cites.
Numeracy for language models: Evaluating and improving their ability to predict numbers. In ACL , Vol. 56. 2104–2115
GP Spithourakis and S Riedel. 2018 · 2018
Earlier work this paper cites.
MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms. In NAACL . 2357–2367
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019 · 2019
Earlier work this paper cites.
Holist: An environment for machine learning of higher order logic theorem proving. In ICML . 454–463
Kshitij Bansal, Sarah Loos, Markus Rabe, Christian Szegedy, and Stewart Wilcox. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs. In NAACL . 2368–2378
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
How Large Are Lions? Inducing Distributions over Quantitative Attributes. In ACL . 3973–3983
Yanai Elazar, Abhijit Mahabal, Deepak Ramachandran, Tania Bedrax-Weiss, and Dan Roth. 2019 · 2019
Earlier work this paper cites.
A comprehensive survey of deep learning for image captioning
MD Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. 2019 · 2019
Earlier work this paper cites.
GamePad: A Learning Environment for Theorem Proving. In ICLR
Daniel Huang, Prafulla Dhariwal, Dawn Song, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
How Much Coffee Was Consumed During EMNLP 2019? Fermi Problems: A New Reasoning Challenge for AI. In EMNLP . 7318–7328
Ashwin Kalyan, Abhinav Kumar, Arjun Chandrasekaran, Ashish Sabharwal, and Peter Clark. 2021 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Metamath: a computer language for mathematical proofs
Norman Megill and David A Wheeler. 2019 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Analysing Mathematical Reasoning Abilities of Neural Models. In ICLR
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019 · 2019
Earlier work this paper cites.
Do NLP Models Know Numbers? Probing Numeracy in Embeddings. In EMNLP . 5307–5315
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Learning to prove theorems via interacting with proof assistants. In ICML . 6984–6994
Kaiyu Yang and Jia Deng. 2019 · 2019
Earlier work this paper cites.
An empirical investigation of contextualized number prediction. In EMNLP . 4754–4764
Taylor Berg-Kirkpatrick and Daniel Spokoyny. 2020 · 2020
Earlier work this paper cites.
Injecting Numerical Reasoning Skills into Language Models. In ACL , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). 946–958
Mor Geva, Ankit Gupta, and Jonathan Berant. 2020 · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Point to the Expression: Solving Algebraic Word Problems using the Expression-Pointer Transformer Model. In EMNLP . Online, 3768–3779
Bugeun Kim, Kyung Seo Ki, Donggeon Lee, and Gahgene Gweon. 2020 · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL . 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers. In ACL . 975–984
Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020 · 2020
Earlier work this paper cites.
Generative Language Modeling for Automated Theorem Proving
Stanislas Polu and Ilya Sutskever. 2020b · 2020
Earlier work this paper cites.
Semantically-Aligned Universal Tree-Structured Solver for Math Word Problems. In EMNLP . 3780–3789
Jinghui Qin, Lihui Lin, Xiaodan Liang, Rumin Zhang, and Liang Lin. 2020 · 2020
Earlier work this paper cites.
Pre-trained models for natural language processing: A survey
Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020 · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
GeoQA: A geometric question answering benchmark towards multimodal numerical reasoning
Jiaqi Chen, Jianheng Tang, Jinghui Qin, Xiaodan Liang, Lingbo Liu, Eric P Xing, and Liang Lin. 2021b · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Advancing mathematics by guiding human intuition with AI
Alex Davies, Petar Veličković, Lars Buesing, Sam Blackwell, Daniel Zheng, Nenad Tomašev, Richard Tanburn, Peter Battaglia, Charles Blundell, András Juhász, et al · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a · 2021
Earlier work this paper cites.
LISA: Language models of ISAbelle proofs. In AITP . 378–392
Albert Qiaochu Jiang, Wenda Li, Jesse Michael Han, and Yuhuai Wu. 2021 · 2021
Earlier work this paper cites.
IsarStep: a Benchmark for High-level Mathematical Reasoning. In ICLR
Wenda Li, Lei Yu, Yuhuai Wu, and Lawrence C Paulson. 2021 · 2021
Earlier work this paper cites.
Investigating the limitations of transformers with simple arithmetic tasks
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin. 2021 · 2021
Earlier work this paper cites.
Pretrained Language Models are Symbolic Mathematics Solvers too!
Kimia Noorbakhsh, Modar Sulaiman, Mahdi Sharifi, Kallol Roy, and Pooyan Jamshidi. 2021 · 2021
Earlier work this paper cites.
Show Your Work: Scratchpads for Intermediate Computation with Language Models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021 · 2021
Earlier work this paper cites.
MathBERT: A Pre-Trained Model for Mathematical Formula Understanding
Shuai Peng, Ke Yuan, Liangcai Gao, and Zhi Tang. 2021 · 2021
Earlier work this paper cites.
Generate & Rank: A Multi-task Framework for Math Word Problems
Jianhao Shen, Yichun Yin, Lin Li, Lifeng Shang, Xin Jiang, Ming Zhang, and Qun Liu. 2021 · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 billion parameter autoregressive language model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Earlier work this paper cites.
Exploring Generalization Ability of Pretrained Language Models on Arithmetic and Logical Reasoning. In NLPCC . 758–769
Cunxiang Wang, Boyuan Zheng, Yuchen Niu, and Yue Zhang. 2021 · 2021
Cited alongside, same era.
NaturalProofs: Mathematical Theorem Proving in Natural Language. In NeurIPS Datasets and Benchmarks Track (Round 1)
Sean Welleck, Jiacheng Liu, Ronan Le Bras, Hannaneh Hajishirzi, Yejin Choi, and Kyunghyun Cho. 2021 · 2021
Cited alongside, same era.
INT: An Inequality Benchmark for Evaluating Generalization in Theorem Proving. In ICLR
Yuhuai Wu, Albert Jiang, Jimmy Ba, and Roger Baker Grosse. 2021 · 2021
Cited alongside, same era.
TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance. In ACL . 3277–3287
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021 · 2021
Cited alongside, same era.
ArMATH: a Dataset for Solving Arabic Math Word Problems. In LREC . 351–362
Reem Alghamdi, Zhenwen Liang, and Xiangliang Zhang. 2022 · 2022
The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language Models. In EMNLP
Jingyuan Qi, Zhiyang Xu, Ying Shen, Minqian Liu, dingnan jin, Qifan Wang, and Lifu Huang. 2023 · 2023
Closest in time.
Creator: Tool creation for disentangling abstract and concrete reasoning of large language models
Cheng Qian, Chi Han, Yi R Fung, Yujia Qin, Zhiyuan Liu, and Heng Ji. 2023 · 2023
Closest in time.
Reasoning with Language Model Prompting: A Survey. In ACL . 5368–5393
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2023 · 2023
Closest in time.
A survey of hallucination in large foundation models
Vipula Rawte, Amit Sheth, and Amitava Das. 2023 · 2023
Closest in time.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022a · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Cited alongside, same era.
A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level
Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, et al · 2022
Cited alongside, same era.
Injecting Numerical Reasoning Skills into Knowledge Base Question Answering Models
Yu Feng, Jing Zhang, Xiaokang Zhang, Lemao Liu, Cuiping Li, and Hong Chen. 2022 · 2022
Cited alongside, same era.
Proof Artifact Co-Training for Theorem Proving with Language Models. In ICLR
Jesse Michael Han, Jason Rute, Yuhuai Wu, Edward Ayers, and Stanislas Polu. 2022 · 2022
Cited alongside, same era.
PGDP5K: A Diagram Parsing Dataset for Plane Geometry Problems
Yihan Hao, Mingliang Zhang, Fei Yin, and Linlin Huang. 2022 · 2022
Cited alongside, same era.
Algorithm of thoughts: Enhancing exploration of ideas in large language models
Bilgehan Sel, Ahmad Al-Tawaha, Vanshaj Khattar, Lu Wang, Ruoxi Jia, and Ming Jin. 2023 · 2023
Closest in time.
Reflexion: Language agents with verbal reinforcement learning. In NeurIPS
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023 · 2023
Closest in time.
SCREWS: A Modular Framework for Reasoning with Revisions
Kumar Shridhar, Harsh Jhamtani, Hao Fang, Benjamin Van Durme, Jason Eisner, and Patrick Xia. 2023 · 2023
Closest in time.
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data. In EMNLP . 12113–12139
Kashun Shum, Shizhe Diao, and Tong Zhang. 2023 · 2023
Closest in time.
AdaPlanner: Adaptive Planning from Feedback with Language Models
Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang. 2023 · 2023
Closest in time.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Improving Mathematics Tutoring With A Code Scratchpad. In BEA . 20–28
Shriyash Upadhyay, Etan Ginsberg, and Chris Callison-Burch. 2023 · 2023
Closest in time.
Making large language models better reasoners with alignment
Peiyi Wang, Lei Li, Liang Chen, Feifan Song, Binghuai Lin, Yunbo Cao, Tianyu Liu, and Zhifang Sui. 2023a · 2023
Closest in time.
Mint: Evaluating llms in multi-turn interaction with tools and language feedback
Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2023c · 2023
Closest in time.
Generative AI for Math: Part I–MathPile: A Billion-Token-Scale Pretraining Corpus for Math
Zengzhi Wang, Rui Xia, and Pengfei Liu. 2023e · 2023
Closest in time.
llmstep: LLM proofstep suggestions in Lean. In NeurIPS
Sean Welleck and Rahul Saha. 2023 · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023 · 2023
Closest in time.
GPT Can Solve Mathematical Problems Without a Calculator
Zhen Yang, Ming Ding, Qingsong Lv, Zhihuan Jiang, Zehai He, Yuyi Guo, Jinfeng Bai, and Jie Tang. 2023 · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023b · 2023
Closest in time.
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Large Language Models
Yao Yao, Zuchao Li, and Hai Zhao. 2023a · 2023
Closest in time.
Answering questions by meta-reasoning over multiple chains of thought
Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, and Jonathan Berant. 2023 · 2023
Closest in time.
Thought Propagation: An Analogical Approach to Complex Reasoning with Large Language Models
Junchi Yu, Ran He, and Rex Ying. 2023a · 2023
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023b · 2023
Closest in time.
Scaling relationship on learning mathematical reasoning with large language models
Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Chuanqi Tan, and Chang Zhou. 2023a · 2023
Closest in time.
How well do Large Language Models perform in Arithmetic tasks?
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang. 2023b · 2023
Closest in time.
Mammoth: Building math generalist models through hybrid instruction tuning
Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023 · 2023
Closest in time.
Interpretable math word problem solution generation via step-by-step planning
Mengxue Zhang, Zichao Wang, Zhichao Yang, Weiqi Feng, and Andrew Lan. 2023b · 2023
Closest in time.
Instruction tuning for large language models: A survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al · 2023
Closest in time.
Cumulative reasoning with large language models
Yifan Zhang, Jingqin Yang, Yang Yuan, and Andrew Chi-Chih Yao. 2023c · 2023
Closest in time.
Verify-and-edit: A knowledge-enhanced chain-of-thought framework
Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023a · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Closest in time.
miniF2F: a cross-system benchmark for formal Olympiad-level mathematics. In ICLR
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. 2023 · 2023
Closest in time.
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Closest in time.
Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification
Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu, Anya Jia, Linqi Song, Mingjie Zhan, et al · 2023
Closest in time.
Language agent tree search unifies reasoning acting and planning in language models
Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2023e · 2023
Closest in time.
ISR-LLM: Iterative Self-Refined Large Language Model for Long-Horizon Sequential Task Planning
Zhehua Zhou, Jiayang Song, Kunpeng Yao, Zhan Shu, and Lei Ma. 2023b · 2023
Closest in time.
Diving into Self-Evolve Training for Multimodal Reasoning. In Submitted to The Thirteenth International Conference on Learning Representations
Anonymous. 2024 · 2024
Closest in time.
Noise contrastive alignment of language models with explicit rewards
Huayu Chen, Guande He, Lifan Yuan, Ganqu Cui, Hang Su, and Jun Zhu. 2024a · 2024
Closest in time.
ControlMath: Controllable Data Generation Promotes Math Generalist Models
Nuo Chen, Ning Wu, Jianhui Chang, and Jia Li. 2024c · 2024
Closest in time.
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
Konstantin Chernyshev, Vitaliy Polshkov, Ekaterina Artemova, Alex Myasnikov, Vlad Stepanov, Alexei Miasnikov, and Sergei Tilga. 2024 · 2024
Closest in time.
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning
Yihe Deng and Paul Mineiro. 2024 · 2024
Closest in time.
Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching. In CIKM . 3730–3735
Yuyang Ding, Hanglei Hu, Jie Zhou, Qin Chen, Bo Jiang, and Liang He. 2024 · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. 2024 · 2024
Closest in time.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al · 2024
Closest in time.
FOLIO: Natural Language Reasoning with First-Order Logic. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024 , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, 22017–22031
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, and Dragomir Radev. 2024 · 2024
Closest in time.
V-star: Training verifiers for self-taught reasoners
Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron Courville, Alessandro Sordoni, and Rishabh Agarwal. 2024 · 2024
Closest in time.
Fewer is More: Boosting Math Reasoning with Reinforced Context Pruning. In EMNLP . 13674–13695
Xijie Huang, Li Lyna Zhang, Kwang-Ting Cheng, Fan Yang, and Mao Yang. 2024 · 2024
Closest in time.
Gpt-4o system card
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Closest in time.
LeanReasoner: Boosting Complex Logical Reasoning with Lean
Dongwei Jiang, Marcio Fonseca, and Shay B Cohen. 2024 · 2024
Closest in time.
Training language models to self-correct via reinforcement learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al · 2024
Closest in time.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia. 2024 · 2024
Closest in time.
Let’s Verify Step by Step. In ICLR
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024 · 2024
Closest in time.
Cmm-math: A chinese multimodal math dataset to evaluate and enhance the mathematics reasoning of large multimodal models
Wentao Liu, Qianjun Pan, Yi Zhang, Zhuo Liu, Ji Wu, Jie Zhou, Aimin Zhou, Qin Chen, Bo Jiang, and Liang He. 2024 · 2024
Closest in time.
Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, et al · 2024
Closest in time.
Reft: Reasoning with reinforced fine-tuning
Trung Quoc Luong, Xinbo Zhang, Zhanming Jie, Peng Sun, Xiaoran Jin, and Hang Li. 2024 · 2024
Closest in time.
OpenAI O1 System Card
OpenAI. 2024 · 2024
Closest in time.
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu, Chong Sun, Xiaoshuai Song, Zhuoma GongQue, Shanglin Lei, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Yifan Zhang, Xiao Zong, Yida Xu, Muxi Diao, Zhimin Bao, Chen Li, and Honggang Zhang. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Wenhao Shi, Zhiqiang Hu, Yi Bin, Junhua Liu, Yang Yang, See-Kiong Ng, Lidong Bing, and Roy Ka-Wei Lee. 2024 · 2024
Closest in time.
Towards large language models as copilots for theorem proving in lean
Peiyang Song, Kaiyu Yang, and Anima Anandkumar. 2024 · 2024
Closest in time.
QwQ: Reflect Deeply on the Boundaries of the Unknown
Qwen Team. 2024 · 2024
Closest in time.
Reft: Reasoning with reinforced fine-tuning. In ACL . 7601–7614
Luong Trung, Xinbo Zhang, Zhanming Jie, Peng Sun, Xiaoran Jin, and Hang Li. 2024 · 2024
Closest in time.
Measuring multimodal mathematical reasoning with math-vision dataset
Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, Mingjie Zhan, and Hongsheng Li. 2024b · 2024
Closest in time.
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024a · 2024
Closest in time.
Understanding, Abstracting and Checking: Evoking Complicated Multimodal Reasoning in LMMs
Yifan wang and Yun Fu. 2024 · 2024
Closest in time.
Evaluating Mathematical Reasoning Beyond Accuracy
Shijie Xia, Xuefeng Li, Yixin Liu, Tongshuang Wu, and Pengfei Liu. 2024 · 2024
Closest in time.
AtomThink: A Slow Thinking Framework for Multimodal Mathematical Reasoning
Kun Xiang, Zhili Liu, Zihao Jiang, Yunshuang Nie, Runhui Huang, Haoxiang Fan, Hanhui Li, Weiran Huang, Yihan Zeng, Jianhua Han, et al · 2024
Closest in time.
LLaVA-o1: Let Vision Language Models Reason Step-by-Step
Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, and Li Yuan. 2024b · 2024
Closest in time.
Faithful Logical Reasoning via Symbolic Chain-of-Thought
Jundong Xu, Hao Fei, Liangming Pan, Qian Liu, Mong-Li Lee, and Wynne Hsu. 2024a · 2024
Closest in time.
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
Yibo Yan, Jiamin Su, Jianxiang He, Fangteng Fu, Xu Zheng, Yuanhuiyi Lyu, Kun Wang, Shen Wang, Qingsong Wen, and Xuming Hu. 2024 · 2024
Closest in time.
Qwen2. 5-math technical report: Toward mathematical expert model via self-improvement
An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, et al · 2024
Closest in time.
MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning
Shuo Yin, Weihao You, Zhilong Ji, Guoqiang Zhong, and Jinfeng Bai. 2024 · 2024
Closest in time.
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
Huaiyuan Ying, Shuo Zhang, Linyang Li, Zhejian Zhou, Yunfan Shao, Zhaoye Fei, Yichuan Ma, Jiawei Hong, Kuikun Liu, Ziyi Wang, Yudong Wang, Zijian Wu, Shuaibin Li, Fengzhe Zhou, Hongwei Liu, Songyang Zhang, Wenwei Zhang, Hang Yan, Xipeng Qiu, Jiayu Wang, Kai Chen, and Dahua Lin. 2024 · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, et al · 2024
Closest in time.
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024 · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D Goodman. 2024 · 2024
Closest in time.
Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b
Di Zhang, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, and Wanli Ouyang. 2024b · 2024
Closest in time.
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
Zhihan Zhang, Tao Ge, Zhenwen Liang, Wenhao Yu, Dian Yu, Mengzhao Jia, Dong Yu, and Meng Jiang. 2024a · 2024
Closest in time.
Marco-o1: Towards open reasoning models for open-ended solutions
Yu Zhao, Huifeng Yin, Bo Zeng, Hao Wang, Tianqi Shi, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. 2024c · 2024
Closest in time.
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
Zilong Zhao, Yao Rong, Dongyang Guo, Emek Gözlüklü, Emir Gülboy, and Enkelejda Kasneci. 2024a · 2024
Closest in time.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?. In ECCV . 169–186
Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Yu Qiao, et al · 2025
Closest in time.
Are NLP Models really able to Solve Simple Math Word Problems?. In NAACL . 2080–2094
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2094
Closest in time.