Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) and Large Reasoning Models (LRMs) are increasingly used for critical tasks, yet they provide no guarantees about the correctness of their solutions.
Large legal fictions: Profiling legal hallucinations in large language models
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho · 1946
Earlier work this paper cites.
Deep blue
Murray Campbell, A. Joseph Hoane, and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
Selecting methods for the analysis of reliance on automation
Lu Wang, Greg Jamieson, and Justin Hollands · 2008
Earlier work this paper cites.
"why should i trust you?": Explaining the predictions of any classifier, 2016
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning, 2017
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (xai)
Amina Adadi and Mohammed Berrada · 2018
Earlier work this paper cites.
On human predictions with explanations and predictions of machine learning models: A case study on deception detection
Vivian Lai and Chenhao Tan · 2019
Earlier work this paper cites.
Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making
Yunfeng Zhang, Q. Vera Liao, and Rachel K. E. Bellamy · 2020
Earlier work this paper cites.
To trust or to think: Cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos · 2021
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Earlier work this paper cites.
Should i follow ai-based advice? measuring appropriate reliance in human-ai decision-making, 2022
Max Schemmer, Patrick Hemmer, Niklas Kühl, Carina Benz, and Gerhard Satzger · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman · 2023
Earlier work this paper cites.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al · 2023
Earlier work this paper cites.
Have llms advanced enough? a challenging problem solving benchmark for large language models, 2023
Daman Arora, Himanshu Gaurav Singh, and Mausam · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Cited alongside, same era.
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Explainable human-ai interaction: A planning perspective, 2024
Gemini Team, Google · 2025
Later among the works it cites.
Claude’s extended thinking
Anthropic · 2025
Later among the works it cites.
gpt-oss-120b & gpt-oss-20b model card, 2025
OpenAI · 2025
Later among the works it cites.
What large language models know and what people think they know
Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas W. Mayer, and Padhraic Smyth · 2025
Later among the works it cites.
Fostering appropriate reliance on large language models: The role of explanations, sources, and inconsistencies
Sunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao, Tania Lombrozo, and Olga Russakovsky · 2025
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarath Sreedharan, Anagha Kulkarni, and Subbarao Kambhampati · 2024
Cited alongside, same era.
Learning to reason with LLMs
OpenAI · 2024
Cited alongside, same era.
Raymond Fok and Daniel S. Weld · 2024
Cited alongside, same era.
Large language models help humans verify truthfulness – except when they are convincingly wrong
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, and Jordan Boyd-Graber · 2024
Cited alongside, same era.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms, 2024
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2024
Cited alongside, same era.
Why would you suggest that? human trust in language model responses, 2024
Manasi Sharma, Ho Chit Siu, Rohan Paleja, and Jaime D. Peña · 2024
Cited alongside, same era.
DeepSeek-V3 technical report, 2024
DeepSeek-AI · 2024
Cited alongside, same era.
Qwen Team · 2024
Cited alongside, same era.
Jessica Y. Bo, Sophia Wan, and Ashton Anderson · 2025
Later among the works it cites.
Qwq-32b: Embracing the power of reinforcement learning, March 2025
Qwen Team · 2025
Later among the works it cites.
Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al · 2025
Later among the works it cites.
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al · 2025
Later among the works it cites.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Later among the works it cites.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars
Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile, and Noah D Goodman · 2025
Later among the works it cites.
Chain-of-thought reasoning in the wild is not always faithful
Iván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, and Arthur Conmy · 2025
Later among the works it cites.
Lujain Ibrahim, Franziska Sofia Hafner, and Luc Rocher · 2025
Later among the works it cites.
A survey of large language models, 2026
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen · 2026
Closest in time.
Siddhant Bhambri, Upasana Biswas, and Subbarao Kambhampati · 2026
Closest in time.
Position: Stop anthropomorphizing intermediate tokens as reasoning/thinking traces!, 2026
Subbarao Kambhampati, Karthik Valmeekam, Siddhant Bhambri, Vardhan Palod, Lucas Saldyt, Kaya Stechly, Soumya Rani Samineni, Durgesh Kalwar, and Upasana Biswas · 2026
Closest in time.
Do llms have core beliefs?
Nitesh Chawla, Anna Sokol, and Marianna Ganapini · 2026
Closest in time.