Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) often fail on complex reasoning tasks due to flawed question comprehension, not just flawed logic.
“Program induction by rationale generation: Learning to solve and explain algebraic word problems,”
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom, · 2017
Earlier work this paper cites.
“Training verifiers to solve math word problems,”
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al., · 2021
Earlier work this paper cites.
“A diverse corpus for evaluating and developing english math word problem solvers,”
Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su, · 2021
Earlier work this paper cites.
“Chain-of-thought prompting elicits reasoning in large language models,”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al., · 2022
Earlier work this paper cites.
“Decomposed prompting: A modular approach for solving complex tasks,”
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal, · 2022
Earlier work this paper cites.
“Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models,”
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim, · 2023
Earlier work this paper cites.
“Echoprompt: instructing the model to rephrase queries for improved in-context learning,”
Rajasekhar Reddy Mekala, Yasaman Razeghi, and Sameer Singh, · 2023
Earlier work this paper cites.
“Rephrase and respond: Let large language models ask better questions for themselves,”
Yihe Deng, Weitong Zhang, Zixiang Chen, and Quanquan Gu, · 2023
Cited alongside, same era.
“Re-reading improves reasoning in large language models,”
Xiaohan Xu, Chongyang Tao, Tao Shen, Can Xu, Hongbo Xu, Guodong Long, Jian-Guang Lou, and Shuai Ma, · 2024
Cited alongside, same era.
“Insights into llm long-context failures: When transformers know but don’t tell,” 2024
Taiming Lu, Muhan Gao, Kuai Yu, Adam Byerly, and Daniel Khashabi, · 2024
Cited alongside, same era.
Qihuang Zhong, Kang Wang, Ziyang Xu, Juhua Liu, Liang Ding, and Bo Du, · 2024
Cited alongside, same era.
“Adapting large language models to domains via reading comprehension,” 2024
Daixuan Cheng, Shaohan Huang, and Furu Wei, · 2024
“Large language models in bioinformatics: A survey,”
Zhenyu Wang, Zikang Wang, Jiyue Jiang, Pengan Chen, Xiangyu Shi, and Yu Li, · 2025
Closest in time.
Feijiang Han, Jiaming Zhang, Chuyi Deng, Jianheng Tang, and Yunhuai Liu, · 2025
Closest in time.
“Logiccat: A chain-of-thought text-to-sql benchmark for multi-domain reasoning challenges,”
Tao Liu, Hongying Zan, Yifan Li, Dixuan Zhang, Lulu Kong, Haixin Liu, Jiaming Hou, Aoze Zheng, Rui Li, Yiming Qiao, et al., · 2025
Closest in time.
“Zerotuning: Unlocking the initial token’s power to enhance large language models without training,”
Feijiang Han, Xiaodong Yu, Jianheng Tang, and Lyle Ungar, · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Bellm: Backward dependency enhanced large language model for sentence embeddings,” 2024
Xianming Li and Jing Li, · 2024
Cited alongside, same era.
“Attributes as textual genes: Leveraging llms as genetic algorithm simulators for conditional synthetic data generation,” 2025
Guangzeng Han, Weisi Liu, and Xiaolei Huang, · 2025
Closest in time.
“Can large language models detect errors in long chain-of-thought reasoning?,”
Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, et al., · 2025
Closest in time.