Fetching the paper…
Reading the bibliography…
Although recent advancements in large language models (LLMs) have significantly improved their performance on various tasks, they still face challenges with complex and symbolic multi-step reasoning, particularly in mathematical reasoning.
Word reordering and a dynamic programming beam search algorithm for statistical machine translation
C. Tillmann and H. Ney · 2003
Earlier work this paper cites.
Multi-armed bandits with episode context
C. D. Rosin · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He · 2021
Earlier work this paper cites.
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Cited alongside, same era.
Model card and evaluations for claude models
Anthropic · 2023
Cited alongside, same era.
Llemma: An open language model for mathematics
Z. Azerbayev, H. Schoelkopf, K. Paster, M. D. Santos, S. McAleer, A. Q. Jiang, J. Deng, S. Biderman, and S. Welleck · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
T. Dao · 2023
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
X. Feng, Z. Wan, M. Wen, Y. Wen, W. Zhang, and J. Wang · 2023
Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning, 2023
K. Wang, H. Ren, A. Zhou, Z. Lu, S. Luo, W. Shi, R. Zhang, L. Song, M. Zhan, and H. Li · 2023
Later among the works it cites.
Large language models are better reasoners with self-verification
Y. Weng, M. Zhu, F. Xia, B. Li, S. He, S. Liu, B. Sun, K. Liu, and J. Zhao · 2023
Later among the works it cites.
Decomposition enhances reasoning via self-evaluation guided decoding, 2023
Y. Xie, K. Kawaguchi, Y. Zhao, X. Zhao, M.-Y. Kan, J. He, and Q. Xie · 2023
Later among the works it cites.
Mammoth: Building math generalist models through hybrid instruction tuning
X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Later among the works it cites.
Solving math word problems via cooperative reasoning induced language models
X. Zhu, J. Wang, L. Zhang, Y. Zhang, Y. Huang, R. Gan, J. Zhang, and Y. Yang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pal: Program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Z. Gou, Z. Shao, Y. Gong, Y. Yang, M. Huang, N. Duan, W. Chen, et al · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Cited alongside, same era.
Query and response augmentation cannot help out-of-domain math reasoning generalization
C. Li, Z. Yuan, G. Dong, K. Lu, J. Wu, C. Tan, X. Wang, and C. Zhou · 2023
Cited alongside, same era.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Cited alongside, same era.
Don’t throw away your value model! making ppo even better via value-guided monte-carlo tree search decoding
J. Liu, A. Cohen, R. Pasunuru, Y. Choi, H. Hajishirzi, and A. Celikyilmaz · 2023
Cited alongside, same era.
SEER: Facilitating structured reasoning and explanation via reinforcement learning
G. Chen, K. Tang, C. Yang, F. Ye, Y. Qiao, and Y. Qian · 2024
Closest in time.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Closest in time.
Mustard: Mastering uniform synthesis of theorem and proof data
Y. Huang, X. Lin, Z. Liu, Q. Cao, H. Xin, H. Wang, Z. Li, L. Song, and X. Liang · 2024
Closest in time.
Mario: Math reasoning with code interpreter output–a reproducible pipeline
M. Liao, W. Luo, C. Li, J. Wu, and K. Fan · 2024
Closest in time.
Z. Lu, A. Zhou, H. Ren, K. Wang, W. Shi, J. Pan, M. Zhan, and H. Li · 2024
Closest in time.
Don’t forget your reward values: Language model alignment via value-based calibration
X. Mao, F.-L. Li, H. Xu, W. Zhang, and A. T. Luu · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. Li, Y. Wu, and D. Guo · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Mario eval: Evaluate your math llm with your math llm–a mathematical dataset evaluation toolkit
B. Zhang, C. Li, and K. Fan · 2024
Closest in time.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Y. Zheng, R. Zhang, J. Zhang, Y. Ye, Z. Luo, and Y. Ma · 2024
Closest in time.