Fetching the paper…
Reading the bibliography…
Large language models (LLMs) often improve their performance in downstream tasks when they generate Chain of Thought reasoning text before producing an answer.
Episodic and semantic memory
Endel Tulving · 1972
Earlier work this paper cites.
From neuropsychology to mental structure
Tim Shallice · 1988
Earlier work this paper cites.
MAWPS: A math word problem repository
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi · 2016
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Earlier work this paper cites.
A diverse corpus for evaluating and developing english math word problem solvers
Shen-yun Miao, Chao-Chun Liang, and Keh-Yih Su · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal · 2021
Earlier work this paper cites.
Why exposure bias matters: An imitation learning perspective of error accumulation in language generation
Kushal Arora, Layla El Asri, Hareesh Bahuleyan, and Jackie Cheung · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Text and patterns: For effective chain of thought, it takes two to tango, 2022
Aman Madaan and Amir Yazdanbakhsh · 2022
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners, 2022
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them, 2022
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei · 2022
Earlier work this paper cites.
Multimodal chain-of-thought reasoning in language models, 2023
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola · 2022
Cited alongside, same era.
Opt-r: Exploring the role of explanations in finetuning and prompting for reasoning skills of large language models
Badr Alkhamissi, Siddharth Verma, Ping Yu, Zhijing Jin, Asli Celikyilmaz, and Mona Diab · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Cited alongside, same era.
Revealing the structure of language model capabilities, 2023
Ryan Burnell, Han Hao, Andrew R. A. Conway, and Jose Hernandez Orallo · 2023
Cited alongside, same era.
Language model behavior: A comprehensive survey, 2023
Tyler A. Chang and Benjamin K. Bergen · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Visual chain of thought: Bridging logical gaps with multimodal infillings, 2023
Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang · 2023
Later among the works it cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting, 2023
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman · 2023
Later among the works it cites.
Towards understanding chain-of-thought prompting: An empirical study of what matters
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Faith and fate: Limits of transformers on compositionality, 2023
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi · 2023
Cited alongside, same era.
Function calling and other api updates, Jul 2023
Atty Eleti, Jeff Harris, and Logan Kilpatrick · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Shapley value attribution in chain of thought, Apr 2023
Leo Gao · 2023
Cited alongside, same era.
An automatically discovered chain-of-thought prompt generalizes to novel models and datasets, 2023
Konstantin Hebenstreit, Robert Praas, Louis P Kiesewetter, and Matthias Samwald · 2023
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning, 2023
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez · 2023
Cited alongside, same era.
Sources of hallucination by large language models on inference tasks, 2023
Nick McKenna, Tianyi Li, Liang Cheng, Mohammad Javad Hosseini, Mark Johnson, and Mark Steedman · 2023
Cited alongside, same era.
Analyzing chain-of-thought prompting in large language models via gradient-based feature attributions, 2023
Skyler Wu, Eric Meng Shen, Charumathi Badrinath, Jiaqi Ma, and Himabindu Lakkaraju · 2023
Later among the works it cites.
Natural language reasoning, a survey, 2023
Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang · 2023
Later among the works it cites.
Faithfulness vs. plausibility: On the (un)reliability of explanations from large language models, 2024
Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju · 2024
Closest in time.
Llms with chain-of-thought are non-causal reasoners, 2024
Guangsheng Bao, Hongbo Zhang, Linyi Yang, Cunxiang Wang, and Yue Zhang · 2024
Closest in time.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning, 2024
Subhabrata Dutta, Joykirat Singh, Soumen Chakrabarti, and Tanmoy Chakraborty · 2024
Closest in time.
Direct evaluation of chain-of-thought in multi-hop reasoning with knowledge graphs, 2024
Minh-Vuong Nguyen, Linhao Luo, Fatemeh Shiri, Dinh Phung, Yuan-Fang Li, Thuy-Trang Vu, and Gholamreza Haffari · 2024
Closest in time.