Fetching the paper…
Reading the bibliography…
Mathematical reasoning and optimization are fundamental to artificial intelligence and computational problem-solving.
H. Gelernter, J. R. Hansen, and D. W. Loveland, “Empirical explorations of the geometry theorem machine,” in western joint IRE-AIEE-ACM computer conference
1960
Earlier work this paper cites.
D. Bobrow et al
1964
Earlier work this paper cites.
L. Nelson, “The socratic method,” vol. 2, pp. 34–38, 1980
1980
Earlier work this paper cites.
D. J. Briars and J. H. Larkin, “An integrated model of skill in solving elementary word problems,” Cognition and instruction
1984
Earlier work this paper cites.
C. R. Fletcher, “Understanding and solving arithmetic word problems: A computer simulation,” Behavior Research Methods, Instruments, & Computers
1985
Earlier work this paper cites.
J. D. Hunter, “Matplotlib: A 2d graphics environment,” Computing in science & engineering
2007
Earlier work this paper cites.
W. McKinney et al
2010
Earlier work this paper cites.
U. Met Office, “Cartopy: A cartographic python library with a matplotlib interface,” Exeter, Devon
2010
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, et al
2011
Earlier work this paper cites.
S. Mitchell, M. OSullivan, and I. Dunning, “Pulp: a linear programming toolkit for python,” The University of Auckland, Auckland, New Zealand
2011
Earlier work this paper cites.
M. J. Hosseini, H. Hajishirzi, O. Etzioni, and N. Kushman, “Learning to solve arithmetic word problems with verb categorization,” in EMNLP
2014
Earlier work this paper cites.
N. Kushman, Y. Artzi, L. Zettlemoyer, and R. Barzilay, “Learning to automatically solve algebra word problems,” in ACL
2014
Earlier work this paper cites.
L. Zhou, S. Dai, and L. Chen, “Learn to solve algebra word problems using quadratic programming,” in EMNLP
2015
Earlier work this paper cites.
S. Roy, T. Vieira, and D. Roth, “Reasoning about quantities in natural language,” TACL
2015
Earlier work this paper cites.
S. Roy and D. Roth, “Solving general arithmetic word problems,” in EMNLP
2015
Earlier work this paper cites.
S. Shi, Y. Wang, C.-Y. Lin, X. Liu, and Y. Rui, “Automatically solving number word problems by semantic parsing and reasoning,” in EMNLP
2015
Earlier work this paper cites.
R. Koncel-Kedziorski, H. Hajishirzi, A. Sabharwal, O. Etzioni, and S. D. Ang, “Parsing algebraic word problems into equations,” TACL
2015
Earlier work this paper cites.
S. Upadhyay and M.-W. Chang, “Draw: A challenging and diverse algebra word problem set,” tech. rep., Citeseer, 2015
2015
Earlier work this paper cites.
A. Grabowski, A. Korniłowicz, and A. Naumowicz, “Four decades of mizar: Foreword,” Journal of Automated Reasoning
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in ICCV
2015
Earlier work this paper cites.
A. Mitra and C. Baral, “Learning to use formulas to solve simple arithmetic problems,” in ACL
2016
Earlier work this paper cites.
G. Irving, C. Szegedy, A. A. Alemi, N. Eén, F. Chollet, and J. Urban, “Deepmath-deep sequence models for premise selection,” NeurIPS
2016
Earlier work this paper cites.
G. Spithourakis, I. Augenstein, and S. Riedel, “Numerically grounded language models for semantic error correction,” in EMNLP
2016
Earlier work this paper cites.
D. Huang, S. Shi, C.-Y. Lin, J. Yin, and W.-Y. Ma, “How well do computers solve math word problems? large-scale dataset construction and evaluation,” in ACL
2016
Earlier work this paper cites.
R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi, “Mawps: A math word problem repository,” in NAACL
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Wang, X. Liu, and S. Shi, “Deep neural solver for math word problems,” in EMNLP
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS
2017
Earlier work this paper cites.
M. Forbes and Y. Choi, “Verb physics: Relative physical knowledge of actions and objects,” in ACL
2017
Earlier work this paper cites.
W. Ling, D. Yogatama, C. Dyer, and P. Blunsom, “Program induction by rationale generation: Learning to solve and explain algebraic word problems,” in ACL
2017
Earlier work this paper cites.
C. Kaliszyk, F. Chollet, and C. Szegedy, “Holstep: A machine learning dataset for higher-order logic theorem proving,” in ICLR
2017
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving Language Understanding by Generative Pre-Training,” p. 12, June 2018
2018
Earlier work this paper cites.
G. Spithourakis and S. Riedel, “Numeracy for language models: Evaluating and improving their ability to predict numbers,” in ACL
2018
Earlier work this paper cites.
S. Roy and D. Roth, “Mapping to declarative knowledge for word problem solving,” TACL
2018
Earlier work this paper cites.
PhD thesis, Universitat Politècnica de València, 2018
S. Alemany Ibor, Diseño e implementación de un simulador basado en agentes estilo JGOMAS en Python · 2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” 2019
2019
Earlier work this paper cites.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach,” 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” p. 24, Feb. 2019
2019
Earlier work this paper cites.
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner, “Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs,” in NAACL
2019
Earlier work this paper cites.
E. Wallace, Y. Wang, S. Li, S. Singh, and M. Gardner, “Do nlp models know numbers? probing numeracy in embeddings,” in EMNLP
2019
Earlier work this paper cites.
K. Yang and J. Deng, “Learning to prove theorems via interacting with proof assistants,” in ICML
2019
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL
2019
Earlier work this paper cites.
Y. Elazar, A. Mahabal, D. Ramachandran, T. Bedrax-Weiss, and D. Roth, “How large are lions? inducing distributions over quantitative attributes,” in ACL
2019
Earlier work this paper cites.
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi, “Mathqa: Towards interpretable math word problem solving with operation-based formalisms,” in NAACL
2019
Earlier work this paper cites.
K. Bansal, S. Loos, M. Rabe, C. Szegedy, and S. Wilcox, “Holist: An environment for machine learning of higher order logic theorem proving,” in ICML
2019
Earlier work this paper cites.
D. Saxton, E. Grefenstette, F. Hill, and P. Kohli, “Analysing mathematical reasoning abilities of neural models,” in ICLR
2019
Earlier work this paper cites.
D. Huang, P. Dhariwal, D. Song, and I. Sutskever, “Gamepad: A learning environment for theorem proving,” in ICLR
2019
Earlier work this paper cites.
M. Z. Hossain, F. Sohel, M. F. Shiratuddin, and H. Laga, “A comprehensive survey of deep learning for image captioning,” ACM Computing Surveys (CsUR)
2019
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in ACL
2020
Earlier work this paper cites.
M. Geva, A. Gupta, and J. Berant, “Injecting numerical reasoning skills into language models,” in ACL
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al
2020
Earlier work this paper cites.
X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, and X. Huang, “Pre-trained models for natural language processing: A survey,” Science China Technological Sciences
2020
Earlier work this paper cites.
X. Zhang, D. Ramachandran, I. Tenney, Y. Elazar, and D. Roth, “Do language embeddings capture scales?,” in BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP
2020
Earlier work this paper cites.
T. Berg-Kirkpatrick and D. Spokoyny, “An empirical investigation of contextualized number prediction,” in EMNLP
2020
Earlier work this paper cites.
S. Polu and I. Sutskever, “Generative language modeling for automated theorem proving,” 2020
2020
Earlier work this paper cites.
J. Maynez, S. Narayan, B. Bohnet, and R. McDonald, “On faithfulness and factuality in abstractive summarization,” in ACL
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” JMLR
2020
Earlier work this paper cites.
B. Kim, K. S. Ki, D. Lee, and G. Gweon, “Point to the Expression: Solving Algebraic Word Problems using the Expression-Pointer Transformer Model,” in EMNLP
2020
Earlier work this paper cites.
J. Qin, L. Lin, X. Liang, R. Zhang, and L. Lin, “Semantically-aligned universal tree-structured solver for math word problems,” in EMNLP
2020
Earlier work this paper cites.
S.-Y. Miao, C.-C. Liang, and K.-Y. Su, “A diverse corpus for evaluating and developing english math word problem solvers,” in ACL
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Berg-Kirkpatrick and D. Spokoyny, “An empirical investigation of contextualized number prediction,” in EMNLP
2020
Earlier work this paper cites.
C. R. Harris, K. J. Millman, S. J. Van Der Walt, R. Gommers, P. Virtanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, et al
2020
Earlier work this paper cites.
P. Clark, O. Etzioni, D. Khashabi, T. Khot, B. D. Mishra, K. Richardson, A. Sabharwal, C. Schoenick, O. Tafjord, N. Tandon, S. Bhakthavatsalam, D. Groeneveld, M. Guerquin, and M. Schmitz, “From ’F’ to ’a’ on the N.Y. regents science exams: An overview of the aristo project,” 2021
2021
Earlier work this paper cites.
S. Peng, K. Yuan, L. Gao, and Z. Tang, “MathBERT: A pre-trained model for mathematical formula understanding,” 2021
2021
Earlier work this paper cites.
A. Q. Jiang, W. Li, J. M. Han, and Y. Wu, “Lisa: Language models of isabelle proofs,” in AITP
2021
Earlier work this paper cites.
F. Zhu, W. Lei, Y. Huang, C. Wang, S. Zhang, J. Lv, F. Feng, and T.-S. Chua, “Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance,” in ACL
2021
Earlier work this paper cites.
R. Nogueira, Z. Jiang, and J. Lin, “Investigating the limitations of transformers with simple arithmetic tasks,” arXiv
2021
Earlier work this paper cites.
C. Wang, B. Zheng, Y. Niu, and Y. Zhang, “Exploring generalization ability of pretrained language models on arithmetic and logical reasoning,” in NLPCC
2021
Earlier work this paper cites.
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, C. Sutton, and A. Odena, “Show Your Work: Scratchpads for Intermediate Computation with Language Models,” 2021
2021
Earlier work this paper cites.
S. Welleck, J. Liu, R. Le Bras, H. Hajishirzi, Y. Choi, and K. Cho, “Naturalproofs: Mathematical theorem proving in natural language,” in NeurIPS Datasets and Benchmarks Track (Round 1)
2021
Earlier work this paper cites.
Y. Wu, A. Jiang, J. Ba, and R. B. Grosse, “Int: An inequality benchmark for evaluating generalization in theorem proving,” in ICLR
2021
Earlier work this paper cites.
K. Noorbakhsh, M. Sulaiman, M. Sharifi, K. Roy, and P. Jamshidi, “Pretrained language models are symbolic mathematics solvers too!,” arXiv
2021
Earlier work this paper cites.
J. Shen, Y. Yin, L. Li, L. Shang, X. Jiang, M. Zhang, and Q. Liu, “Generate & Rank: A Multi-task Framework for Math Word Problems,” 2021
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” ICLR
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the math dataset,” in NeurIPS Datasets and Benchmarks Track (Round 2)
2021
Earlier work this paper cites.
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, “Training verifiers to solve math word problems,” 2021
2021
Earlier work this paper cites.
A. Patel, S. Bhattamishra, and N. Goyal, “Are nlp models really able to solve simple math word problems?,” in NAACL
2021
Earlier work this paper cites.
A. Kalyan, A. Kumar, A. Chandrasekaran, A. Sabharwal, and P. Clark, “How much coffee was consumed during emnlp 2019? fermi problems: A new reasoning challenge for ai,” in EMNLP
2021
Earlier work this paper cites.
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. R. Routledge, et al
2021
Earlier work this paper cites.
W. Li, L. Yu, Y. Wu, and L. C. Paulson, “Isarstep: a benchmark for high-level mathematical reasoning,” in ICLR
2021
Earlier work this paper cites.
J. Chen, J. Tang, J. Qin, X. Liang, L. Liu, E. P. Xing, and L. Lin, “Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning,” arXiv
2021
Earlier work this paper cites.
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. Kim, J. Kim, Y. Tae, C. Park, J.-H. Choi, and J. Choo, “Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift,” in ICLR
2021
Earlier work this paper cites.
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, et al
2021
Earlier work this paper cites.
P. Lu, L. Qiu, J. Chen, T. Xia, Y. Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu, “Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning,” in NeurIPS Datasets and Benchmarks Track
2021
Earlier work this paper cites.
B. Wang and A. Komatsuzaki, “Gpt-j-6b: A 6 billion parameter autoregressive language model,” 2021
2021
Earlier work this paper cites.
A. Davies, P. Veličković, L. Buesing, S. Blackwell, D. Zheng, N. Tomašev, R. Tanburn, P. Battaglia, C. Blundell, A. Juhász, et al
2021
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al
2022
Earlier work this paper cites.
Y. Feng, J. Zhang, X. Zhang, L. Liu, C. Li, and H. Chen, “Injecting numerical reasoning skills into knowledge base question answering models,” 2022
2022
Earlier work this paper cites.
Y. Zhao, Y. Li, C. Li, and R. Zhang, “Multihiertt: Numerical reasoning over multi hierarchical tabular and textual data,” in ACL
2022
Earlier work this paper cites.
Z. Jie, J. Li, and W. Lu, “Learning to reason deductively: Math word problem solving as complex relation extraction,” 2022
2022
Earlier work this paper cites.
Z. Li, W. Zhang, C. Yan, Q. Zhou, C. Li, H. Liu, and Y. Cao, “Seeking patterns, not just memorizing procedures: Contrastive learning for solving math word problems,” 2022
2022
Earlier work this paper cites.
S. Min, M. Lewis, L. Zettlemoyer, and H. Hajishirzi, “Metaicl: Learning to learn in context,” in NAACL
2022
Cited alongside, same era.
Y. Chen, R. Zhong, S. Zha, G. Karypis, and H. He, “Meta-learning via language model in-context tuning,” in ACL
2022
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al
2022
Cited alongside, same era.
W. Chen, X. Ma, X. Wang, and W. W. Cohen, “Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks,” arXiv
2022
Cited alongside, same era.
I. Drori, S. Zhang, R. Shuttleworth, L. Tang, A. Lu, E. Ke, K. Liu, L. Chen, S. Tran, N. Cheng, et al
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Saravia, A. Poulton, V. Kerkez, and R. Stojnic, “Galactica: A Large Language Model for Science,” 2022
2022
Cited alongside, same era.
H. Zhou, A. Nova, H. Larochelle, A. Courville, B. Neyshabur, and H. Sedghi, “Teaching algorithmic reasoning via in-context learning,” 2022
2022
Cited alongside, same era.
A. Q. Jiang, W. Li, S. Tworkowski, K. Czechowski, T. Odrzygóźdź, P. Miłoś, Y. Wu, and M. Jamnik, “Thor: Wielding hammers to integrate language models and automated theorem provers,” 2022
2022
Cited alongside, same era.
J. M. Han, J. Rute, Y. Wu, E. Ayers, and S. Polu, “Proof artifact co-training for theorem proving with language models,” in ICLR
2022
Cited alongside, same era.
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al
2022
Cited alongside, same era.
Z. Liang, J. Zhang, L. Wang, W. Qin, Y. Lan, J. Shao, and X. Zhang, “MWP-BERT: Numeracy-augmented pre-training for math word problem solving,” 2022
2022
Cited alongside, same era.
M. Kazemi, N. Kim, D. Bhatia, X. Xu, and D. Ramachandran, “Lambada: Backward chaining for automated reasoning in natural language,” arXiv
2022
Cited alongside, same era.
2023
Later among the works it cites.
H. He, H. Zhang, and D. Roth, “Socreval: Large language models with the socratic method for reference-free reasoning evaluation,” arXiv
2023
Later among the works it cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al
2023
Later among the works it cites.
A. Garza and M. Mergenthaler-Canseco, “TimeGPT-1,” arXiv preprint arXiv:2310.03589
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Li, C. Liu, S. Cheng, R. Arcucci, and S. Hong, “Frozen Language Model Helps ECG Zero-Shot Learning,” in MIDL
2023
Later among the works it cites.
L. Y. Jiang, X. C. Liu, N. P. Nejatian, M. Nasir-Moin, D. Wang, A. Abidin, et al
2023
Later among the works it cites.
X. Wang, M. Fang, Z. Zeng, and T. Cheng, “Where Would I Go Next? Large Language Models as Human Mobility Predictors,” 2023
2023
Later among the works it cites.
H. Xue and F. D. Salim, “PromptCast: A New Prompt-based Learning Paradigm for Time Series Forecasting,” IEEE TKDE
2023
Later among the works it cites.
N. Gruver, M. Finzi, S. Qiu, and A. G. Wilson, “Large Language Models Are Zero Shot Time Series Forecasters,” in NeurIPS
2023
Later among the works it cites.
T. Zhou, P. Niu, X. Wang, L. Sun, and R. Jin, “One Fits All: Power General Time Series Analysis by Pretrained LM,” in NeurIPS
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Shi, S. Xue, K. Wang, F. Zhou, J. Y. Zhang, J. Zhou, C. Tan, and H. Mei, “Language Models Can Improve Event Prediction by Few-Shot Abductive Reasoning,” in NeurIPS
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Rasul, A. Ashok, A. R. Williams, A. Khorasani, G. Adamopoulos, R. Bhagwatkar, et al
2023
Later among the works it cites.
C. Mai, Y. Chang, C. Chen, and Z. Zheng, “Enhanced Scalable Graph Neural Network via Knowledge Distillation,” IEEE TNNLS
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
V. Rawte, A. Sheth, and A. Das, “A survey of hallucination in large foundation models,” arXiv
2023
Later among the works it cites.
A. Zhou, K. Wang, Z. Lu, W. Shi, S. Luo, Z. Qin, S. Lu, A. Jia, L. Song, M. Zhan, et al
2023
Later among the works it cites.
Y. Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen, “Making language models better reasoners with step-aware verifier,” in ACL
2023
Later among the works it cites.
R. Zhao, X. Li, S. Joty, C. Qin, and L. Bing, “Verify-and-edit: A knowledge-enhanced chain-of-thought framework,” arXiv
2023
Later among the works it cites.
K. Shridhar, H. Jhamtani, H. Fang, B. Van Durme, J. Eisner, and P. Xia, “Screws: A modular framework for reasoning with revisions,” arXiv
2023
Later among the works it cites.
J. Gawlikowski, C. R. N. Tassi, M. Ali, J. Lee, M. Humt, J. Feng, A. Kruspe, R. Triebel, P. Jung, R. Roscher, et al
2023
Later among the works it cites.
J. Duan, H. Cheng, S. Wang, C. Wang, A. Zavalny, R. Xu, B. Kailkhura, and K. Xu, “Shifting attention to relevance: Towards the uncertainty estimation of large language models,” arXiv
2023
Later among the works it cites.
W. Zhou, Y. E. Jiang, E. Wilcox, R. Cotterell, and M. Sachan, “Controlled text generation with natural language instructions,” in ICML
2023
Later among the works it cites.
A. N. Bernardino Romera-Paredes, Mohammadamin Barekatain et al
2023
Later among the works it cites.
N. Ding, Y. Chen, B. Xu, Y. Qin, S. Hu, Z. Liu, M. Sun, and B. Zhou, “Enhancing chat language models by scaling high-quality instructional conversations,” in EMNLP
2023
Later among the works it cites.
Y. Yan, J. Su, J. He, F. Fu, X. Zheng, Y. Lyu, K. Wang, S. Wang, Q. Wen, and X. Hu, “A survey of mathematical reasoning in the era of multimodal large language model: Benchmark, method & challenges,” arXiv
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Zhang, T. Ge, Z. Liang, W. Yu, D. Yu, M. Jia, D. Yu, and M. Jiang, “Learn beyond the answer: Training language models with reflection for mathematical reasoning,” arXiv
2024
Later among the works it cites.
L. Yuan, G. Cui, H. Wang, N. Ding, X. Wang, J. Deng, B. Shan, H. Chen, R. Xie, Y. Lin, et al
2024
Later among the works it cites.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” NeurIPS
2024
Later among the works it cites.
H. Chen, G. He, L. Yuan, G. Cui, H. Su, and J. Zhu, “Noise contrastive alignment of language models with explicit rewards,” arXiv
2024
Later among the works it cites.
A. Yang, B. Zhang, B. Hui, B. Gao, B. Yu, C. Li, D. Liu, J. Tu, J. Zhou, J. Lin, et al
2024
Later among the works it cites.
H. Ying, S. Zhang, L. Li, Z. Zhou, Y. Shao, Z. Fei, Y. Ma, J. Hong, K. Liu, Z. Wang, Y. Wang, Z. Wu, S. Li, F. Zhou, H. Liu, S. Zhang, W. Zhang, H. Yan, X. Qiu, J. Wang, K. Chen, and D. Lin, “Internlm-math: Open math large language models toward verifiable reasoning,” 2024
2024
Later among the works it cites.
D. Jiang, M. Fonseca, and S. B. Cohen, “Leanreasoner: Boosting complex logical reasoning with lean,” arXiv
2024
Later among the works it cites.
Z. Zhao, Y. Rong, D. Guo, E. Gözlüklü, E. Gülboy, and E. Kasneci, “Stepwise self-consistent mathematical reasoning with large language models,” arXiv
2024
Later among the works it cites.
L. Luo, Y. Liu, R. Liu, S. Phatale, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun, et al
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
K. M. Collins, A. Q. Jiang, S. Frieder, L. Wong, M. Zilka, U. Bhatt, T. Lukasiewicz, Y. Wu, J. B. Tenenbaum, W. Hart, et al
2024
Later among the works it cites.
N. Chen, N. Wu, J. Chang, and J. Li, “Controlmath: Controllable data generation promotes math generalist models,” arXiv
2024
Later among the works it cites.
S. Yin, W. You, Z. Ji, G. Zhong, and J. Bai, “Mumath-code: Combining tool-use large language models with multi-perspective data augmentation for mathematical reasoning,” arXiv
2024
Later among the works it cites.
A. Hosseini, X. Yuan, N. Malkin, A. Courville, A. Sordoni, and R. Agarwal, “V-star: Training verifiers for self-taught reasoners,” arXiv
2024
Later among the works it cites.
E. Zelikman, G. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. D. Goodman, “Quiet-star: Language models can teach themselves to think before speaking,” arXiv
2024
Later among the works it cites.
T. Q. Luong, X. Zhang, Z. Jie, P. Sun, X. Jin, and H. Li, “Reft: Reasoning with reinforced fine-tuning,” arXiv
2024
Later among the works it cites.
A. Kumar, V. Zhuang, R. Agarwal, Y. Su, J. D. Co-Reyes, A. Singh, K. Baumli, S. Iqbal, C. Bishop, R. Roelofs, et al
2024
Later among the works it cites.
D. Zhang, X. Huang, D. Zhou, Y. Li, and W. Ouyang, “Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b,” arXiv
2024
Later among the works it cites.
Y. Zhao, H. Yin, B. Zeng, H. Wang, T. Shi, C. Lyu, L. Wang, W. Luo, and K. Zhang, “Marco-o1: Towards open reasoning models for open-ended solutions,” arXiv
2024
Later among the works it cites.
X. Lai, Z. Tian, Y. Chen, S. Yang, X. Peng, and J. Jia, “Step-dpo: Step-wise preference optimization for long-chain reasoning of llms,” arXiv
2024
Later among the works it cites.
Y. Deng and P. Mineiro, “Flow-dpo: Improving llm mathematical reasoning through online multi-agent learning,” arXiv
2024
Later among the works it cites.
Y. Ding, H. Hu, J. Zhou, Q. Chen, B. Jiang, and L. He, “Boosting large language models with socratic method for conversational mathematics teaching,” in CIKM
2024
Later among the works it cites.
Q. Team, “Qwq: Reflect deeply on the boundaries of the unknown,” 2024
2024
Later among the works it cites.
OpenAI, “OpenAI O1 System Card,” 2024
2024
Later among the works it cites.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al
2024
Later among the works it cites.
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y. Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin, “Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution,” arXiv
2024
Later among the works it cites.
G. Xu, P. Jin, L. Hao, Y. Song, L. Sun, and L. Yuan, “Llava-o1: Let vision language models reason step-by-step,” arXiv
2024
Later among the works it cites.
T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, et al
2024
Later among the works it cites.
K. Xiang, Z. Liu, Z. Jiang, Y. Nie, R. Huang, H. Fan, H. Li, W. Huang, Y. Zeng, J. Han, et al
2024
Later among the works it cites.
W. Shi, Z. Hu, Y. Bin, J. Liu, Y. Yang, S.-K. Ng, L. Bing, and R. K.-W. Lee, “Math-llava: Bootstrapping mathematical reasoning for multimodal large language models,” 2024
2024
Later among the works it cites.
Anonymous, “Diving into self-evolve training for multimodal reasoning,” in Submitted to The Thirteenth International Conference on Learning Representations
2024
Later among the works it cites.
Y. wang and Y. Fu, “Understanding, abstracting and checking: Evoking complicated multimodal reasoning in LMMs,” 2024
2024
Later among the works it cites.
J. Zhao, J. Tong, Y. Mou, M. Zhang, Q. Zhang, and X.-J. Huang, “Exploring the compositional deficiency of large language models in mathematical reasoning through trap problems,” in EMNLP
2024
Later among the works it cites.
K. Wang, J. Pan, W. Shi, Z. Lu, M. Zhan, and H. Li, “Measuring multimodal mathematical reasoning with math-vision dataset,” arXiv
2024
Later among the works it cites.
W. Liu, Q. Pan, Y. Zhang, Z. Liu, J. Wu, J. Zhou, A. Zhou, Q. Chen, B. Jiang, and L. He, “Cmm-math: A chinese multimodal math dataset to evaluate and enhance the mathematics reasoning of large multimodal models,” arXiv
2024
Later among the works it cites.
Q. Chen, L. Qin, J. Zhang, Z. Chen, X. Xu, and W. Che, “M 3 CoT: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought,” in ACL
2024
Later among the works it cites.
S. Xia, X. Li, Y. Liu, T. Wu, and P. Liu, “Evaluating mathematical reasoning beyond accuracy,” arXiv
2024
Later among the works it cites.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen, “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” 2024
2024
Later among the works it cites.
R. Qiao, Q. Tan, G. Dong, M. Wu, C. Sun, X. Song, Z. GongQue, S. Lei, Z. Wei, M. Zhang, R. Qiao, Y. Zhang, X. Zong, Y. Xu, M. Diao, Z. Bao, C. Li, and H. Zhang, “We-math: Does your large multimodal model achieve human-like mathematical reasoning?,” 2024
2024
Later among the works it cites.
K. Chernyshev, V. Polshkov, E. Artemova, A. Myasnikov, V. Stepanov, A. Miasnikov, and S. Tilga, “U-math: A university-level benchmark for evaluating mathematical skills in llms,” 2024
2024
Later among the works it cites.
S. Han, H. Schoelkopf, Y. Zhao, Z. Qi, M. Riddell, W. Zhou, J. Coady, D. Peng, Y. Qiao, L. Benson, L. Sun, A. Wardle-Solano, H. Szabó, E. Zubova, M. Burtell, J. Fan, Y. Liu, B. Wong, M. Sailor, A. Ni, L. Nan, J. Kasai, T. Yu, R. Zhang, A. R. Fabbri, W. Kryscinski, S. Yavuz, Y. Liu, X. V. Lin, S. Joty, Y. Zhou, C. Xiong, R. Ying, A. Cohan, and D. Radev, “FOLIO: natural language reasoning with first-order logic,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024
2024
Later among the works it cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” in ICLR
2024
Later among the works it cites.
T. H. Trinh, Y. Wu, Q. V. Le, H. He, and T. Luong, “Solving olympiad geometry without human demonstrations,” Nature
2024
Later among the works it cites.
DeepSeek-AI, “Deepseek LLM: scaling open-source language models with longtermism,” CoRR
2024
Later among the works it cites.
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang, “Deepseek-coder: When the large language model meets programming – the rise of code intelligence,” 2024
2024
Later among the works it cites.
C. Liu, S. Yang, Q. Xu, Z. Li, C. Long, Z. Li, and R. Zhao, “Spatial-Temporal Large Language Model for Traffic Prediction,” 2024
2024
Later among the works it cites.
D. Cao, F. Jia, S. O. Arik, T. Pfister, Y. Zheng, W. Ye, and Y. Liu, “TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting,” in ICLR
2024
Later among the works it cites.
C. Sun, Y. Li, H. Li, and S. Hong, “TEST: Text Prototype Aligned Embedding to Activate LLM’s Ability for Time Series,” in ICLR
2024
Later among the works it cites.
M. Jin, S. Wang, L. Ma, Z. Chu, J. Y. Zhang, X. Shi, P.-Y. Chen, Y. Liang, Y.-F. Li, S. Pan, et al
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Anthropic, “Introducing the next generation of claude,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
Z. Li, Z. Zhou, Y. Yao, X. Zhang, Y.-F. Li, C. Cao, F. Yang, and X. Ma, “Neuro-symbolic data generation for math reasoning,” Advances in Neural Information Processing Systems
2025
Closest in time.
R. Zhang, D. Jiang, Y. Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K.-W. Chang, Y. Qiao, et al
2025
Closest in time.
2025
Closest in time.