Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown excellent mastering of human language, but still struggle in real-world applications that require mathematical problem-solving.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Mawps: A math word problem repository
R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
W. Ling, D. Yogatama, C. Dyer, and P. Blunsom · 2017
Earlier work this paper cites.
TL;DR: Mining Reddit to learn automatic summarization
M. Völske, M. Potthast, S. Syed, and B. Stein · 2017
Earlier work this paper cites.
Deep neural solver for math word problems
Y. Wang, X. Liu, and S. Shi · 2017
Earlier work this paper cites.
Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al · 2019
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
D. Saxton, E. Grefenstette, F. Hill, and P. Kohli · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. Le bras, J. Gao, and Y. Choi · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Generative language modeling for automated theorem proving
S. Polu and I. Sutskever · 2020
Earlier work this paper cites.
Ape210k: A large-scale and template-rich dataset of math word problems, 2020
W. Zhao, M. Shang, Y. Liu, L. Wang, and J. Liu · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment, 2021
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, J. Kernion, K. Ndousse, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, and J. Kaplan · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Adversarial examples for evaluating math word problem solvers
V. Kumar, R. Maheshwary, and V. Pudi · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang · 2022
Earlier work this paper cites.
CSL: A large-scale Chinese scientific literature dataset
Y. Li, Y. Zhang, Z. Zhao, L. Shen, W. Liu, W. Mao, and H. Zhang · 2022
Earlier work this paper cites.
Numglue: A suite of fundamental yet challenging mathematical reasoning tasks
S. Mishra, A. Mitra, N. Varshney, B. Sachdeva, P. Clark, C. Baral, and A. Kalyan · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, et al · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, et al · 2022
Cited alongside, same era.
Introducing claude, 2023
Anthropic · 2023
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations, 2023
P. Wang, L. Li, Z. Shao, R. X. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models, 2023
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Qwen technical report, 2023
J. Bai, S. Bai, et al · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, et al · 2023
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, L. Gianinazzi, J. Gajda, T. Lehmann, M. Podstawski, H. Niewiadomski, P. Nyczyk, et al · 2023
Cited alongside, same era.
Generative ai for math: Abel
E. Chern, H. Zou, X. Li, J. Hu, K. Feng, J. Li, and P. Liu · 2023
Cited alongside, same era.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Y. Fu, L. Ou, M. Chen, Y. Wan, H. Peng, and T. Khot · 2023
Cited alongside, same era.
Tora: A tool-integrated reasoning agent for mathematical problem solving, 2023
Z. Gou, Z. Shao, Y. Gong, Y. Shen, Y. Yang, M. Huang, N. Duan, and W. Chen · 2023
Cited alongside, same era.
P. Ke, B. Wen, Z. Feng, X. Liu, X. Lei, J. Cheng, S. Wang, A. Zeng, Y. Dong, H. Wang, et al · 2023
Cited alongside, same era.
Cmath: Can your language model pass chinese elementary school math test?, 2023
T. Wei, J. Luan, W. Liu, S. Dong, and B. Wang · 2023
Later among the works it cites.
Large language models as optimizers
C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou, and X. Chen · 2023
Later among the works it cites.
Skymath: Technical report, 2023
L. Yang, H. Yang, W. Cheng, L. Lin, C. Li, Y. Chen, L. Liu, J. Pan, T. Wei, B. Li, L. Zhao, L. Wang, B. Zhu, G. Li, X. Wu, X. Luo, and R. Hu · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Later among the works it cites.
A series of large language models trained from scratch by developers at 01-ai
Yi · 2023
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
L. Yu, W. Jiang, H. Shi, J. Yu, Z. Liu, Y. Zhang, J. T. Kwok, Z. Li, A. Weller, and W. Liu · 2023
Later among the works it cites.
Scaling relationship on learning mathematical reasoning with large language models, 2023
Z. Yuan, H. Yuan, C. Li, G. Dong, K. Lu, C. Tan, C. Zhou, and J. Zhou · 2023
Later among the works it cites.
How well do large language models perform in arithmetic tasks?, 2023
Z. Yuan, H. Yuan, C. Tan, W. Wang, and S. Huang · 2023
Later among the works it cites.
Mammoth: Building math generalist models through hybrid instruction tuning
X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Later among the works it cites.
Evaluating the performance of large language models on gaokao benchmark
X. Zhang, C. Li, Y. Zong, Z. Ying, L. He, and X. Qiu · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Later among the works it cites.
Characterglm: Customizing chinese conversational ai characters with large language models
J. Zhou, Z. Chen, D. Wan, B. Wen, Y. Song, J. Yu, Y. Huang, L. Peng, J. Yang, X. Xiao, et al · 2023
Later among the works it cites.
Mathattack: Attacking large language models towards math solving ability
Z. Zhou, Q. Wang, M. Jin, J. Yao, J. Ye, W. Liu, W. Wang, X. Huang, and K. Huang · 2023
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism, 2024
DeepSeek-AI, :, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, H. Gao, K. Gao, W. Gao, R. Ge, K. Guan, D. Guo, J. Guo, G. Hao, Z. Hao, Y. He, W. Hu, P. Huang, E. Li, G. Li, J. Li, Y. Li, Y. K. Li, W. Liang, F. Lin, A. X. Liu, B. Liu, W. Liu, X. Liu, X. Liu, Y. Liu, H. Lu, S. Lu, F. Luo, S. Ma, X. Nie, T. Pei, Y. Piao, J. Qiu, H. Qu, T. Ren, Z. Ren, C. Ruan, Z. Sha, Z. Shao, J. Song, X. Su, J. Sun, Y. Sun, M. Tang, B. Wang, P. Wang, S. Wang, Y. Wang, Y. Wang, T. Wu, Y. Wu, X. Xie, Z. Xie, Z. Xie, Y. Xiong, H. Xu, R. X. Xu, Y. Xu, D. Yang, Y. You, S. Yu, X. Yu, B. Zhang, H. Zhang, L. Zhang, L. Zhang, M. Zhang, M. Zhang, W. Zhang, Y. Zhang, C. Zhao, Y. Zhao, S. Zhou, S. Zhou, Q. Zhu, and Y. Zou · 2024
Closest in time.
Chatglm-rlhf: Practices of aligning large language models with human feedback, 2024
Z. Hou, Y. Niu, Z. Du, X. Zhang, X. Liu, A. Zeng, Q. Zheng, M. Huang, H. Wang, J. Tang, and Y. Dong · 2024
Closest in time.
Charactereval: A chinese benchmark for role-playing conversational agent evaluation, 2024
Q. Tu, S. Fan, Z. Tian, and R. Yan · 2024
Closest in time.
Sciglm: Training scientific language models with self-reflective instruction annotation and tuning, 2024
D. Zhang, Z. Hu, S. Zhoubian, Z. Du, K. Yang, Z. Wang, Y. Yue, Y. Dong, and J. Tang · 2024
Closest in time.