Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) demonstrate remarkable capabilities in various reasoning tasks.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
KILT: a benchmark for knowledge intensive language tasks
Petroni, F.; Piktus, A.; Fan, A.; Lewis, P.; Yazdani, M.; De Cao, N.; Thorne, J.; Jernite, Y.; Karpukhin, V.; Maillard, J.; et al. 2020 · 2009
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; et al. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Chen, W.; Ma, X.; Wang, X.; and Cohen, W. W. 2022 · 2022
Earlier work this paper cites.
Complexity-based prompting for multi-step reasoning
Fu, Y.; Peng, H.; Sabharwal, A.; Clark, P.; and Khot, T. 2022 · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Automatic chain of thought prompting in large language models
Zhang, Z.; Zhang, A.; Li, M.; and Smola, A. 2022 · 2022
Earlier work this paper cites.
Have llms advanced enough? a challenging problem solving benchmark for large language models
Arora, D.; Singh, H. G.; et al. 2023 · 2023
Earlier work this paper cites.
Generative ai for math: Abel
Chern, E.; Zou, H.; Li, X.; Hu, J.; Feng, K.; Li, J.; and Liu, P. 2023 · 2023
Cited alongside, same era.
Pal: Program-aided language models
Gao, L.; Madaan, A.; Zhou, S.; Alon, U.; Liu, P.; Yang, Y.; Callan, J.; and Neubig, G. 2023 · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H.; Sun, Q.; Xu, C.; Zhao, P.; Lou, J.; Tao, C.; Geng, X.; Lin, Q.; Chen, S.; and Zhang, D. 2023 · 2023
Cited alongside, same era.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
Miao, N.; Teh, Y. W.; and Rainforth, T. 2023 · 2023
Cited alongside, same era.
Rethinking language models as symbolic knowledge graphs
Mruthyunjaya, V.; Pezeshkpour, P.; Hruschka, E.; and Bhutani, N. 2023 · 2023
Cited alongside, same era.
GeoVQA: A Comprehensive Multimodal Geometry Dataset for Secondary Education
Anand, A.; Jaiswal, R.; Dharmadhikari, A.; Marathe, A.; Popat, H.; Mital, H.; Nair, A. R.; Prasad, K.; Kumar, S.; Verma, A.; et al. 2024b · 2024
Closest in time.
From local to global: A graph rag approach to query-focused summarization
Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; and Larson, J. 2024 · 2024
Closest in time.
He, C.; Luo, R.; Bai, Y.; Hu, S.; Thai, Z. L.; Shen, J.; Hu, J.; Han, X.; Huang, Y.; Zhang, Y.; et al. 2024 · 2024
Closest in time.
Li, X.; Wang, W.; Li, M.; Guo, J.; Zhang, Y.; and Feng, F. 2024 · 2024
Closest in time.
Deductive verification of chain-of-thought reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Structured chemistry reasoning with large language models
Ouyang, S.; Zhang, Z.; Yan, B.; Liu, X.; Han, J.; and Qin, L. 2023 · 2023
Cited alongside, same era.
Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning
Wang, K.; Ren, H.; Zhou, A.; Lu, Z.; Luo, S.; Shi, W.; Zhang, R.; Song, L.; Zhan, M.; and Li, H. 2023 · 2023
Cited alongside, same era.
Expertprompting: Instructing large language models to be distinguished experts
Xu, B.; Yang, A.; Lin, J.; Wang, Q.; Zhou, C.; Zhang, Y.; and Mao, Z. 2023 · 2023
Cited alongside, same era.
Large language models as analogical reasoners
Yasunaga, M.; Chen, X.; Li, Y.; Pasupat, P.; Leskovec, J.; Liang, P.; Chi, E. H.; and Zhou, D. 2023 · 2023
Cited alongside, same era.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L.; Jiang, W.; Shi, H.; Yu, J.; Liu, Z.; Zhang, Y.; Kwok, J. T.; Li, Z.; Weller, A.; and Liu, W. 2023 · 2023
Cited alongside, same era.
Zhou, A.; Wang, K.; Lu, Z.; Shi, W.; Luo, S.; Qin, Z.; Lu, S.; Jia, A.; Song, L.; Zhan, M.; et al. 2023 · 2023
Cited alongside, same era.
Revolutionizing High School Physics Education: A Novel Dataset
Anand, A.; Addala, K.; Baghel, K.; Goel, A.; Hira, M.; Gupta, R.; and Shah, R. R. 2023a
Cited in the paper.
Ling, Z.; Fang, Y.; Li, X.; Huang, Z.; Lee, M.; Memisevic, R.; and Su, H. 2024 · 2024
Closest in time.
SciAgent: Tool-augmented Language Models for Scientific Reasoning
Ma, Y.; Gou, Z.; Hao, J.; Xu, R.; Wang, S.; Pan, L.; Yang, Y.; Cao, Y.; and Sun, A. 2024 · 2024
Closest in time.
Scieval: A multi-level large language model evaluation benchmark for scientific research
Sun, L.; Han, Y.; Zhao, Z.; Ma, D.; Shen, Z.; Chen, B.; Chen, L.; and Yu, K. 2024 · 2024
Closest in time.
LLMs cannot find reasoning errors, but can correct them given the error location
Tyen, G.; Mansoor, H.; Carbune, V.; Chen, P.; and Mak, T. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2024 · 2024
Closest in time.
Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Yuan, Z.; Yuan, H.; Li, C.; Dong, G.; Lu, K.; Tan, C.; Zhou, C.; and Zhou, J. 2024 · 2024
Closest in time.