Fetching the paper…
Reading the bibliography…
This study introduces a hypothesis-testing framework to assess whether large language models (LLMs) possess genuine reasoning abilities or primarily depend on token bias.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Note on the sampling error of the difference between correlated proportions or percentages
Quinn McNemar. 1947 · 1947
Earlier work this paper cites.
Some statistical problems in measuring the subjective response to drugs
Frederick Mosteller. 1952 · 1952
Earlier work this paper cites.
The framing of decisions and the psychology of choice
Amos Tversky and Daniel Kahneman. 1981 · 1981
Earlier work this paper cites.
Extensional versus intuitive reasoning: The conjunction fallacy in probability judgment
Amos Tversky and Daniel Kahneman. 1983 · 1983
Earlier work this paper cites.
Don’t make your llm an evaluation benchmark cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. 2023 · 1983
Earlier work this paper cites.
Testing statistical hypotheses , volume 3
Erich Leo Lehmann, Joseph P Romano, and George Casella. 1986 · 1986
Earlier work this paper cites.
Rational choice and the framing of decisions
Amos Tversky and Daniel Kahneman. 1988 · 1988
Earlier work this paper cites.
Controlling the false discovery rate: A practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg. 1995 · 1995
Earlier work this paper cites.
Rational choice in an uncertain world: The psychology of judgment and decision making
Reid Hastie and Robyn M Dawes. 2009 · 2009
Earlier work this paper cites.
Temporal reasoning on implicit events from distant supervision
Ben Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, and Dan Roth. 2020 · 2010
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman. 2011 · 2011
Earlier work this paper cites.
Categorical data analysis , volume 792
Alan Agresti. 2012 · 2012
Earlier work this paper cites.
A corpus and evaluation framework for deeper understanding of commonsense stories
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016 · 2016
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J Liu, and Christopher D Manning. 2017 · 2017
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al. 2020 · 2020
Earlier work this paper cites.
Language models show human-like content effects on reasoning
Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. 2022 · 2022
Earlier work this paper cites.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang. 2022 · 2022
Earlier work this paper cites.
Zhijing Jin, Abhinav Lalwani, Tejas Vaidhya, Xiaoyu Shen, Yiwen Ding, Zhiheng Lyu, Mrinmaya Sachan, Rada Mihalcea, and Bernhard Schoelkopf. 2022 · 2022
Earlier work this paper cites.
Z-icl: zero-shot in-context learning with pseudo-demonstrations
Xinxi Lyu, Sewon Min, Iz Beltagy, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022 · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Reasoning with language model prompting: A survey
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2022 · 2022
Earlier work this paper cites.
Towards understanding chain-of-thought prompting: An empirical study of what matters
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun. 2022 · 2022
Earlier work this paper cites.
Learning to decompose: Hypothetical question decomposition based on comparable texts
Ben Zhou, Kyle Richardson, Xiaodong Yu, and Dan Roth. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz. 2023 · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Cited alongside, same era.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot. 2023 · 2023
Cited alongside, same era.
Rationality of thought improves reasoning in large language models
Tian Gou, Boyao Zhang, Zhenglie Sun, Jing Wang, Yangang Wang, and Jue Wang. 2023 · 2023
Cited alongside, same era.
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt
Thilo Hagendorff, Sarah Fabi, and Michal Kosinski. 2023 · 2023
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. 2024 · 2024
Closest in time.
T2 of thoughts: Temperature tree elicits reasoning in large language models
Chengkun Cai, Xu Zhao, Yucheng Du, Haoliang Liu, and Lei Li. 2024 · 2024
Closest in time.
Premise order matters in reasoning with large language models
Xinyun Chen, Ryan A Chi, Xuezhi Wang, and Denny Zhou. 2024 · 2024
Closest in time.
Cognitive bias in high-stakes decision-making with llms
Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He. 2024 · 2024
Closest in time.
Puzzle solving using reasoning of large language models: A survey
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. 2023 · 2023
Cited alongside, same era.
A closer look at the self-verification abilities of large language models in logical reasoning
Ruixin Hong, Hongming Zhang, Xinyu Pang, Dong Yu, and Changshui Zhang. 2023 · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Unveiling theory of mind in large language models: A parallel to single neurons in the human brain
Mohsen Jamali, Ziv M Williams, and Jing Cai. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Cited alongside, same era.
Evaluating large language models in theory of mind tasks
Michal Kosinski. 2023 · 2023
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. 2023 · 2023
Cited alongside, same era.
Panagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, and Giorgos Stamou. 2024 · 2024
Closest in time.
The impact of reasoning step length on large language models
Mingyu Jin, Qinkai Yu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, Mengnan Du, et al. 2024 · 2024
Closest in time.
Training language models to self-correct via reinforcement learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al. 2024 · 2024
Closest in time.
Internal consistency and self-feedback in large language models: A survey
Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang, Qingchen Yu, Xunkai Li, Rong-Hua Li, Feiyu Xiong, and Zhiyu Li. 2024 · 2024
Closest in time.
(ir) rationality and cognitive biases in large language models
Olivia Macmillan-Scott and Mirco Musolesi. 2024 · 2024
Closest in time.
A logic for expressing log-precision transformers
William Merrill and Ashish Sabharwal. 2024 · 2024
Closest in time.
Heuristic reasoning in ai: Instrumental use and mimetic absorption
Anirban Mukherjee and Hannah Hanwen Chang. 2024 · 2024
Closest in time.
Agent q: Advanced reasoning and learning for autonomous ai agents
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. 2024 · 2024
Closest in time.
How much are llms contaminated? a comprehensive survey and the llmsanitize library
Mathieu Ravaut, Bosheng Ding, Fangkai Jiao, Hailin Chen, Xingxuan Li, Ruochen Zhao, Chengwei Qin, Caiming Xiong, and Shafiq Joty. 2024 · 2024
Closest in time.
Times man of the year list
Jennifer Rosenberg. 2021 · 2024
Closest in time.
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett. 2024 · 2024
Closest in time.
Do large language models show decision heuristics similar to humans? a case study using gpt-3.5
Gaurav Suri, Lily R Slater, Ali Ziaee, and Morgan Nguyen. 2024 · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.
Occupational employment and wage statistics
USDL. 2024 · 2024
Closest in time.
Pengda Wang, Zilin Xiao, Hanjie Chen, and Frederick L Oswald. 2024 · 2024
Closest in time.
Forbes celebrity 100
Wikipedia contributors. 2024a · 2024
Closest in time.
News media in the united states
Wikipedia contributors. 2024b · 2024
Closest in time.
Logicvista: Multimodal llm logical reasoning benchmark in visual contexts
Yijia Xiao, Edward Sun, Tianyu Liu, and Wei Wang. 2024 · 2024
Closest in time.
Wildfiregpt: Tailored large language model for wildfire analysis
Yangxinyu Xie, Tanwi Mallick, Joshua David Bergerson, John K Hutchison, Duane R Verner, Jordan Branham, M Ross Alexander, Robert B Ross, Yan Feng, Leslie-Anne Levy, et al. 2024 · 2024
Closest in time.
Can speculative sampling accelerate react without compromising reasoning quality?
Han Xu, Jingyang Ye, Yutong Li, and Haipeng Chen. 2024 · 2024
Closest in time.
Do large language models latently perform multi-hop reasoning?
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Attention heads of large language models: A survey
Zifan Zheng, Yezhaohui Wang, Yuxin Huang, Shichao Song, Bo Tang, Feiyu Xiong, and Zhiyu Li. 2024 · 2024
Closest in time.