Fetching the paper…
Reading the bibliography…
While the inconsistency of LLMs is not a novel topic, prior research has predominantly addressed two types of generative inconsistencies: i) Randomness Inconsistency: running the same LLM multiple trials, yielding varying responses; ii) Paraphrase Inconsistency: paraphrased prompts result in different responses from the same LLM.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 1905
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models, 2021
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Condaqa: A contrastive reading comprehension dataset for reasoning about negation
Abhilasha Ravichander, Matt Gardner, and Ana Marasović · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
The falcon series of open language models, 2023
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Earlier work this paper cites.
Robustness of learning from task instructions, 2023
Jiasheng Gu, Hongyu Zhao, Hanzi Xu, Liangyu Nie, Hongyuan Mei, and Wenpeng Yin · 2023
Earlier work this paper cites.
Consistency analysis of chatgpt, 2023
Myeongjun Erik Jang and Thomas Lukasiewicz · 2023
Earlier work this paper cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Cited alongside, same era.
Assessing hidden risks of llms: An empirical study on robustness, consistency, and credibility, 2023
Wentao Ye, Mingfeng Ou, Tianyi Li, Yipeng chen, Xuetao Ma, Yifan Yanggong, Sai Wu, Jie Fu, Gang Chen, Haobo Wang, and Junbo Zhao · 2023
Cited alongside, same era.
LLM stability: A detailed analysis with some surprises
Berk Atil, Alexa Chittams, Liseng Fu, Ferhan Ture, Lixinyu Xu, and Breck Baldwin · 2024
Cited alongside, same era.
One vs. many: Comprehending accurate information from multiple erroneous and inconsistent AI generations
Yoonjoo Lee, Kihoon Son, Tae Soo Kim, Jisu Kim, John Joon Young Chung, Eytan Adar, and Juho Kim · 2024
Later among the works it cites.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment, 2024
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2024
Later among the works it cites.
Aaar-1.0: Assessing ai’s potential to assist research
Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Jian Xie, Yuxuan Sun, Yusen Zhang, Jihyun Janice Ahn, et al · 2024
Later among the works it cites.
Order-independence without fine tuning
Reid McIlroy-Young, Katrina Brown, Conlan Olson, Linjun Zhang, and Cynthia Dwork · 2024
Later among the works it cites.
Build the future of ai with meta llama 3, 2024
Meta · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What did I do wrong? quantifying llms’ sensitivity and consistency to prompt engineering
Federico Errica, Giuseppe Siracusano, Davide Sanvito, and Roberto Bifulco · 2024
Cited alongside, same era.
Assessment and mitigation of inconsistencies in llm-based evaluations
Sarik Ghazarian, Yidong Zou, Swair Shah, Nanyun Peng, Anurag Beniwal, Christopher Potts, and Narayanan Sadagopan · 2024
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al
Cited in the paper.
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
Reasoning aware self-consistency: Leveraging reasoning paths for efficient llm sampling, 2025
Guangya Wan, Yuqi Wu, Jie Chen, and Sheng Li · 2025
Closest in time.
Julian Junyan Wang and Victor Xiaoqi Wang · 2025
Closest in time.