Fetching the paper…
Reading the bibliography…
In reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion.
Fuzzy logic
Lotfi A Zadeh. 2008 · 2008
Earlier work this paper cites.
Deductive reasoning
Phil Johnson-Laird. 2010 · 2010
Earlier work this paper cites.
Quantifying uncertainty in answers from any language model and enhancing their trustworthiness
Jiuhai Chen and Jonas Mueller. 2023 · 2023
Earlier work this paper cites.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023 · 2023
Earlier work this paper cites.
ROSCOE: A suite of metrics for scoring step-by-step reasoning
Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023 · 2023
Earlier work this paper cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023 · 2023
Earlier work this paper cites.
Safe planning in dynamic environments using conformal prediction
Lars Lindemann, Matthew Cleaveland, Gihyun Shim, and George J. Pappas. 2023 · 2023
Earlier work this paper cites.
Cpsign-conformal prediction for cheminformatics modeling
Staffan Arvidsson McShane, Ulf Norinder, Jonathan Alvarsson, Ernst Ahlberg, Lars Carlsson, and Ola Spjuth. 2023 · 2023
Earlier work this paper cites.
ReCEval: Evaluating reasoning chains via correctness and informativeness
Archiki Prasad, Swarnadeep Saha, Xiang Zhou, and Mohit Bansal. 2023 · 2023
Earlier work this paper cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023 · 2023
Earlier work this paper cites.
Faithfulness vs. plausibility: On the (un) reliability of explanations from large language models
Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. 2024 · 2024
Earlier work this paper cites.
Empirical validation of conformal prediction for trustworthy skin lesions classification
Jamil Fayyad, Shadi Alijani, and Homayoun Najjaran. 2024 · 2024
Cited alongside, same era.
Socreval: Large language models with the socratic method for reference-free reasoning evaluation
Hangfeng He, Hongming Zhang, and Dan Roth. 2024 · 2024
Cited alongside, same era.
A chain-of-thought is as strong as its weakest link: A benchmark for verifiers of reasoning chains
Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, and Mor Geva. 2024 · 2024
Cited alongside, same era.
Towards faithful model explanation in NLP: A survey
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch. 2024 · 2024
Cited alongside, same era.
Captaincook4d: A dataset for understanding errors in procedural activities
Rohith Peddi, Shivvrat Arya, Bharath Challa, Likhitha Pallapothula, Akshay Vyas, Bhavya Gouripeddi, Qifan Zhang, Jikai Wang, Vasundhara Komaragiri, Eric Ragan, Nicholas Ruozzi, Yu Xiang, and Vibhav Gogate. 2024 · 2024
BIRD: A trustworthy bayesian inference framework for large language models
Yu Feng, Ben Zhou, Weidong Lin, and Dan Roth. 2025 · 2025
Closest in time.
Can large language models detect errors in long chain-of-thought reasoning?
Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, and Bo Zheng. 2025 · 2025
Closest in time.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and 1 others. 2025 · 2025
Closest in time.
Probabilistic stability guarantees for feature attributions
Helen Jin, Anton Xue, Weiqiu You, Surbhi Goel, and Eric Wong. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How much do prompting methods help llms on quantitative reasoning with irrelevant information?
Seok Hwan Song and Wallapak Tavanapong. 2024 · 2024
Cited alongside, same era.
Step-by-step reasoning to solve grid puzzles: Where do LLMs falter?
Nemika Tyagi, Mihir Parmar, Mohith Kulkarni, Aswin Rrv, Nisarg Patel, Mutsumi Nakamura, Arindam Mitra, and Chitta Baral. 2024 · 2024
Cited alongside, same era.
How easily do irrelevant inputs skew the responses of large language models?
Siye Wu, Jian Xie, Jiangjie Chen, Tinghui Zhu, Kai Zhang, and Yanghua Xiao. 2024 · 2024
Cited alongside, same era.
Natural language reasoning, a survey
Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang. 2024 · 2024
Cited alongside, same era.
Processbench: Identifying process errors in mathematical reasoning
Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2024 · 2024
Cited alongside, same era.
Jinu Lee and Julia Hockenmaier. 2025 · 2025
Closest in time.
Premise-augmented reasoning chains improve error identification in math reasoning with llms
Sagnik Mukherjee, Abhinav Chinta, Takyoung Kim, Tarun Anoop Sharma, and Dilek Hakkani-Tür. 2025 · 2025
Closest in time.
Gpt-4o mini: advancing cost-efficient intelligence
OpenAI. 2024 · 2025
Closest in time.
Prmbench: A fine-grained and challenging benchmark for process-level reward models
Mingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou, and Yu Cheng. 2025 · 2025
Closest in time.
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025 · 2025
Closest in time.
The lessons of developing process reward models in mathematical reasoning
Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2025 · 2025
Closest in time.