Fetching the paper…
Reading the bibliography…
In a plethora of recent work, large language models (LLMs) demonstrated impressive reasoning ability, but many proposed downstream reasoning tasks only focus on final answers.
The fallacy of begging the question
John A Barker. 1976 · 1976
Earlier work this paper cites.
Finding contradictions in text
Marie-Catherine de Marneffe, Anna N. Rafferty, and Christopher D. Manning. 2008 · 2008
Earlier work this paper cites.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasović, and Noah A Smith. 2020 · 2010
Earlier work this paper cites.
Pearson’s correlation coefficient
Philip Sedgwick. 2012 · 2012
Earlier work this paper cites.
Commonsense for generative multi-hop question answering tasks
Lisa Bauer, Yicheng Wang, and Mohit Bansal. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Earlier work this paper cites.
Logical fallacies
Domina Petric. 2020 · 2020
Earlier work this paper cites.
Wikicontradiction: Detecting self-contradiction articles on wikipedia
Cheng Hsu, Cheng-Te Li, Diego Saez-Trumper, and Yi-Zhan Hsu. 2021 · 2021
Earlier work this paper cites.
Few-shot self-rationalization with natural language prompts
Ana Marasović, Iz Beltagy, Doug Downey, and Matthew E Peters. 2021 · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI. 2022 · 2022
Cited alongside, same era.
Does self-rationalization improve robustness to spurious correlations?
Alexis Ross, Matthew E Peters, and Ana Marasović. 2022 · 2022
Cited alongside, same era.
Entailer: Answering questions with faithful and truthful chains of reasoning
Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
The unreliability of explanations in few-shot prompting for textual reasoning
Xi Ye and Greg Durrett. 2022 · 2022
Cited alongside, same era.
Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation
Niels Mündler, Jingxuan He, Slobodan Jenko, and Martin Vechev. 2023 · 2023
Closest in time.
Automated evaluation of written discourse coherence using GPT-4
Ben Naismith, Phoebe Mulcaire, and Jill Burstein. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Tailoring self-rationalizers with multi-reward distillation
Sahana Ramnath, Brihi Joshi, Skyler Hallinan, Ximing Lu, Liunian Harold Li, Aaron Chan, Jack Hessel, Yejin Choi, and Xiang Ren. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Cited alongside, same era.
Reconcile: Round-table conference improves reasoning via consensus among diverse llms
Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal. 2023 · 2023
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023 · 2023
Cited alongside, same era.
Roscoe: A suite of metrics for scoring step-by-step reasoning
Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023 · 2023
Cited alongside, same era.
Gpt-4 as an effective zero-shot evaluator for scientific figure captions
Ting-Yao Hsu, Chieh-Yang Huang, Ryan Rossi, Sungchul Kim, C. Lee Giles, and Ting-Hao K. Huang. 2023 · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023 · 2023
Closest in time.
Self-evaluation guided beam search for reasoning
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie. 2023 · 2023
Closest in time.
Why does chatgpt fall short in providing truthful answers
Shen Zheng, Jie Huang, and Kevin Chen-Chuan Chang. 2023 · 2023
Closest in time.
Chain-of-questions training with latent answers for robust multistep question answering
Wang Zhu, Jesse Thomason, and Robin Jia. 2023 · 2023
Closest in time.
Introducing the next generation of claude
Anthropic. 2024 · 2024
Closest in time.
Shibo Hao, Yi Gu, Haotian Luo, Tianyang Liu, Xiyan Shao, Xinyuan Wang, Shuhua Xie, Haodi Ma, Adithya Samavedhi, Qiyue Gao, Zhen Wang, and Zhiting Hu. 2024 · 2024
Closest in time.
A chain-of-thought is as strong as its weakest link: A benchmark for verifiers of reasoning chains
Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, and Mor Geva. 2024 · 2024
Closest in time.
Are you sure? challenging llms leads to performance drops in the flipflop experiment
Philippe Laban, Lidiya Murakhovs’ka, Caiming Xiong, and Chien-Sheng Wu. 2024 · 2024
Closest in time.
Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation
Niels Mündler, Jingxuan He, Slobodan Jenko, and Martin Vechev. 2024 · 2024
Closest in time.
Soumya Sanyal, Tianyi Xiao, Jiacheng Liu, Wenya Wang, and Xiang Ren. 2024 · 2024
Closest in time.