Fetching the paper…
Reading the bibliography…
Large Reasoning Models (LRMs) significantly improve the reasoning ability of Large Language Models (LLMs) by learning to reason, exhibiting promising performance in solving complex tasks.
W. Ling, D. Yogatama, C. Dyer, and P. Blunsom, “Program induction by rationale generation: Learning to solve and explain algebraic word problems,” in Proc. of ACL , 2017, pp. 158–167
2017
Earlier work this paper cites.
P. Jansen, E. Wainwright, S. Marmorstein, and C. Morrison, “WorldTree: A corpus of explanation graphs for elementary science questions supporting multi-hop inference,” in Proc. of LREC , 2018
2018
Earlier work this paper cites.
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal, “Can a suit of armor conduct electricity? a new dataset for open book question answering,” in EMNLP , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Khot, P. Clark, M. Guerquin, P. A. Jansen, and A. Sabharwal, “QASC: A dataset for question answering via sentence composition,” in AAAI , 2019
2019
Earlier work this paper cites.
A. Talmor, J. Herzig, N. Lourie, and J. Berant, “Commonsenseqa: A question answering challenge targeting commonsense knowledge,” in Proc. of NAACL , 2019
2019
Earlier work this paper cites.
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu, “PubMedQA: A dataset for biomedical research question answering,” in Proc. of EMNLP , 2019, pp. 2567–2577
2019
Earlier work this paper cites.
S.-y. Miao, C.-C. Liang, and K.-Y. Su, “A diverse corpus for evaluating and developing English math word problem solvers,” in Proc. of ACL , 2020, pp. 975–984
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
W. Yu, Z. Jiang, Y. Dong, and J. Feng, “Reclor: A reading comprehension dataset requiring logical reasoning,” in Proc. of ICLR , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the math dataset,” NeurIPS , 2021
2021
Earlier work this paper cites.
S. Aggarwal, D. Mandowara, V. Agrawal, D. Khandelwal, P. Singla, and D. Garg, “Explanations for CommonsenseQA: New Dataset and Models,” in Proc. of ACL , 2021
2021
Earlier work this paper cites.
M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant, “Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies,” Transactions of the Association for Computational Linguistics (TACL) , 2021
2021
Earlier work this paper cites.
T. R. Besold, A. d’Avila Garcez, S. Bader, H. Bowman, P. Domingos, P. Hitzler, K.-U. Kühnberger, L. C. Lamb, P. M. V. Lima, L. de Penning et al. , “Neural-symbolic learning and reasoning: A survey and interpretation 1,” in Neuro-Symbolic Artificial Intelligence: The State of the Art , 2021, pp. 1–51
2021
Earlier work this paper cites.
OpenAI, “Introducing chatgpt,” https://openai.com/index/chatgpt/ , 2022
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Proc. of NeurIPS , pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “Musique: Multihop questions via single-hop question composition,” TACL , pp. 539–554, 2022
2022
Earlier work this paper cites.
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in Proc. of NeurIPS , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Chen, M. Yin, M. Ku, P. Lu, Y. Wan, X. Ma, J. Xu, X. Wang, and T. Xia, “TheoremQA: A theorem-driven question answering dataset,” in Proc. of EMNLP , 2023, pp. 7889–7901
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Saparov and H. He, “Language models are greedy reasoners: A systematic formal analysis of chain-of-thought,” in Proc. of ICLR , 2023
2023
Earlier work this paper cites.
N. L. Rane, A. Tawde, S. P. Choudhary, and J. Rane, “Contribution and performance of chatgpt and other large language models (llm) for scientific and research advancements: a double-edged sword,” International Research Journal of Modernization in Engineering Technology and Science , pp. 875–899, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
T. Liu, Q. Guo, X. Hu, C. Jiayang, Y. Zhang, X. Qiu, and Z. Zhang, “Can language models learn to skip steps?” in Proc. of NeurIPS , 2024
2024
Earlier work this paper cites.
Z. Hu, L. Song, J. Zhang, Z. Xiao, J. Wang, Z. Chen, J. Zhao, and H. Xiong, “Rethinking llm-based preference evaluation,” arXiv e-prints , pp. arXiv–2407, 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
C.-H. Chiang and H.-y. Lee, “Over-reasoning and redundant calculation of large language models,” in Proc. of EACL , 2024, pp. 161–169
2024
Earlier work this paper cites.
H. Liu, Z. Zheng, Y. Qiao, H. Duan, Z. Fei, F. Zhou, W. Zhang, S. Zhang, D. Lin, and K. Chen, “MathBench: Evaluating the theory and application proficiency of LLMs with a hierarchical mathematics benchmark,” in Proc. of ACL Findings , 2024
2024
Earlier work this paper cites.
C. He, R. Luo, Y. Bai, S. Hu, Z. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, J. Liu, L. Qi, Z. Liu, and M. Sun, “OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems,” in Proc. of ACL , 2024, pp. 3828–3850
2024
Earlier work this paper cites.
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman, “GPQA: A graduate-level google-proof q&a benchmark,” in First Conference on Language Modeling , 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun et al. , “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” in Proc. of CVPR , 2024, pp. 9556–9567
2024
Earlier work this paper cites.
X. Wang, Z. Hu, P. Lu, Y. Zhu, J. Zhang, S. Subramaniam, A. R. Loomba, S. Zhang, Y. Sun, and W. Wang, “Scibench: Evaluating college-level scientific problem-solving abilities of large language models,” in Proc. of ICML , 2024
2024
Earlier work this paper cites.
K. Wang, J. Pan, W. Shi, Z. Lu, H. Ren, A. Zhou, M. Zhan, and H. Li, “Measuring multimodal mathematical reasoning with math-vision dataset,” in Proc. of NeurIPS , 2024
2024
Earlier work this paper cites.
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao, “Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,” in Proc. of ICLR , 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
E. Ullah, A. Parwani, M. M. Baig, and R. Singh, “Challenges and barriers of using large language models (llm) such as chatgpt for diagnostic medicine with a focus on digital pathology–a recent scoping review,” Diagnostic pathology , p. 43, 2024
2024
Earlier work this paper cites.
I. Cheong, K. Xia, K. K. Feng, Q. Z. Chen, and A. X. Zhang, “(a) i am not a lawyer, but…: engaging legal experts towards responsible llm policies for legal advice,” in Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , 2024, pp. 2454–2469
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
I. M. Olympiad, “American invitational mathematics examination,” https://artofproblemsolving.com/wiki/index.php/American_Invitational_Mathematics_Examination?srsltid=AfmBOoqo573PtuNmYWTobFVQWyhhDjV2VXowjsIZ0kvmHQ_UP_Jn2wrG/ , 2025
2025
Cited alongside, same era.
OpenAI, “Introducing deep research,” https://openai.com/index/introducing-deep-research/ , 2025
2025
Cited alongside, same era.
OpenAI, “Openai o3-mini system card,” 2025
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Jang, J. Kim, W. Kweon, S. Lee, and H. Yu, “Verbosity-aware rationale reduction: Sentence-level rationale reduction for efficient and effective reasoning,” in Proc. of ACL Findings , 2025, pp. 20 769–20 784
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Ahuja, P. Vaddamanu, and B. Patra, “Efficientxlang: Towards improving token efficiency through cross-lingual reasoning,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
H. Liu, L. Cao, Y. Ren, M. Zhou, H. Dong, X. Ma, S. Han, and D. Zhang, “Bingo: Boosting efficient reasoning of llms via dynamic and significance-based reinforcement learning,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. Ding, Y. Chen, F. Wang, L. Ming, and T. Lin, “Do thinking tokens help or trap? towards more efficient large reasoning model,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
OpenAI, “Detecting misbehavior in frontier reasoning models,” https://openai.com/index/chain-of-thought-monitoring/ , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Thapa, S. Shiwakoti, S. B. Shah, S. Adhikari, H. Veeramani, M. Nasim, and U. Naseem, “Large language models (llm) in computational social science: prospects, current state, and challenges,” Social Network Analysis and Mining , pp. 1–30, 2025
2025
Closest in time.
2025
Closest in time.
OpenAI, “Introducing gpt-4.5,” https://openai.com/index/introducing-gpt-4-5/ , 2025
2025
Closest in time.
2025
Closest in time.
Google, “Gemini robotics brings ai into the physical world,” https://deepmind.google/discover/blog/gemini-robotics-brings-ai-into-the-physical-world/ , 2025
2025
Closest in time.
Nvidia, “Nvidia isaac gr00t n1: An open foundation model for humanoid robots,” https://research.nvidia.com/publication/2025-03_nvidia-isaac-gr00t-n1-open-foundation-model-humanoid-robots , 2025
2025
Closest in time.
2025
Closest in time.
I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs from preference data,” in Proc. of ICLR , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
A. Patel, S. Bhattamishra, and N. Goyal, “Are NLP models really able to solve simple math word problems?” in Proc. of NAACL , 2021, pp. 2080–2094
2094
Closest in time.