Fetching the paper…
Reading the bibliography…
Chain-of-thought emerges as a promising technique for eliciting reasoning capabilities from Large Language Models (LLMs).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Note on the sampling error of the difference between correlated proportions or percentages
Quinn McNemar. 1947 · 1947
Earlier work this paper cites.
Estimating causal effects of treatments in randomized and nonrandomized studies
Donald B Rubin. 1974 · 1974
Earlier work this paper cites.
Causal analysis
David R Heise. 1975 · 1975
Earlier work this paper cites.
Identification and estimation of local average treatment effects
Joshua Angrist and Guido Imbens. 1995 · 1995
Earlier work this paper cites.
Naive theories and causal deduction
Denise Dellarosa Cummins. 1995 · 1995
Earlier work this paper cites.
Mechanical reasoning by mental simulation
Mary Hegarty. 2004 · 2004
Earlier work this paper cites.
Causal reasoning through intervention
York Hagmayer, Steven A Sloman, David A Lagnado, and Michael R Waldmann. 2007 · 2007
Earlier work this paper cites.
Causality
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Proofwriter: Generating implications, proofs, and abductive statements over natural language
Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark. 2020 · 2012
Earlier work this paper cites.
Causal inference in statistics, social, and biomedical sciences
Guido W Imbens and Donald B Rubin. 2015 · 2015
Earlier work this paper cites.
Causality in thought
Steven A Sloman and David Lagnado. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Causal parrots: Large language models may talk causality but are not causal
Matej Zečević, Moritz Willig, Devendra Singh Dhami, and Kristian Kersting. 2023 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
Atomic: An atlas of machine commonsense for if-then reasoning
Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A Smith, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Representation learning via invariant causal mechanisms
Jovana Mitrovic, Brian McWilliams, Jacob C Walker, Lars Holger Buesing, and Charles Blundell. 2020 · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Cited alongside, same era.
Counterfactual invariance to spurious correlations: Why and how to pass stress tests
Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein. 2021 · 2021
Cited alongside, same era.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
T Wu, M Tulio Ribeiro, J Heer, and D Weld. 2021 · 2021
Cited alongside, same era.
Exploring the efficacy of automatically generated counterfactuals for sentiment analysis
Linyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang, Barry Smyth, and Ruihai Dong. 2021 · 2021
Faithful chain-of-thought reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning
Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang. 2023 · 2023
Later among the works it cites.
Reinforcement learning from human feedback: Progress and challenges
John Schulman. 2023 · 2023
Later among the works it cites.
Musr: Testing the limits of chain-of-thought with multistep soft reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022 · 2022
Cited alongside, same era.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Amir Feder, Katherine A Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E Roberts, et al. 2022 · 2022
Cited alongside, same era.
Folio: Natural language reasoning with first-order logic
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, et al. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
Limitations of language models in arithmetic and symbolic induction
Jing Qian, Hong Wang, Zekun Li, Shiyang Li, and Xifeng Yan. 2022 · 2022
Cited alongside, same era.
Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al. 2023 · 2023
Later among the works it cites.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023 · 2023
Later among the works it cites.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023 · 2023
Later among the works it cites.
Fangzhi Xu, Qika Lin, Jiawei Han, Tianzhe Zhao, Jun Liu, and Erik Cambria. 2023 · 2023
Later among the works it cites.
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023 · 2023
Later among the works it cites.
Chain-of-thought unfaithfulness as disguised accuracy
Oliver Bentham, Nathan Stringham, and Ana Marasović. 2024 · 2024
Closest in time.
On the hardness of faithful chain-of-thought reasoning in large language models
Sree Harsha Tanneru, Dan Ley, Chirag Agarwal, and Himabindu Lakkaraju. 2024 · 2024
Closest in time.
The impact of reasoning step length on large language models
Mingyu Jin, Qinkai Yu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, Mengnan Du, et al. 2024 · 2024
Closest in time.
Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024 · 2024
Closest in time.
Let’s think dot by dot: Hidden computation in transformer language models
Jacob Pfau, William Merrill, and Samuel R Bowman. 2024 · 2024
Closest in time.
General purpose verification for chain of thought prompting
Robert Vacareanu, Anurag Pratik, Evangelia Spiliopoulou, Zheng Qi, Giovanni Paolini, Neha Anna John, Jie Ma, Yassine Benajiba, and Miguel Ballesteros. 2024 · 2024
Closest in time.
Dissociation of faithful and unfaithful reasoning in llms
Evelyn Yee, Alice Li, Chenyu Tang, Yeon Ho Jung, Ramamohan Paturi, and Leon Bergen. 2024 · 2024
Closest in time.
Natural language reasoning, a survey
Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang. 2024 · 2024
Closest in time.