Fetching the paper…
Reading the bibliography…
Since the advent of Large Language Models (LLMs), efforts have largely focused on improving their instruction-following and deductive reasoning abilities, leaving open the question of whether these models can truly discover new knowledge.
Abductive commonsense reasoning, 2020
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen tau Yih, and Yejin Choi · 1908
Earlier work this paper cites.
On the measure of intelligence, 2019
François Chollet · 1911
Earlier work this paper cites.
Peirce’s theory of abduction
Arthur W. Burks · 1946
Earlier work this paper cites.
Peirce’s notion of abduction
Harry G. Frankfurt · 1958
Earlier work this paper cites.
The inference to the best explanation
Gilbert H. Harman · 1965
Earlier work this paper cites.
William whewell on the consilience of inductions
Larry Laudan · 1971
Earlier work this paper cites.
Collected papers of charles sanders peirce , volume 5
Charles Sanders Peirce · 1974
Earlier work this paper cites.
A logic for default reasoning
Raymond Reiter · 1980
Earlier work this paper cites.
Some philosophical problems from the standpoint of artificial intelligence
John McCarthy and Patrick J Hayes · 1981
Earlier work this paper cites.
Nonmonotonic logic and temporal projection
Steve Hanks and Drew McDermott · 1987
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Peirce-suit of truth –why inference to the best explanation and abduction ought not to be confused
Gerhard Minnameier · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Deception detection
Pär Anders Granhag and Aldert Vrij · 2005
Earlier work this paper cites.
The logic of scientific discovery
Karl Popper · 2005
Earlier work this paper cites.
The road to experience and prediction from within: Hans reichenbach’s scientific correspondence from berlin to istanbul
Friedrich Stadler · 2011
Earlier work this paper cites.
The effect of wording on message propagation: Topic- and author-controlled natural experiments on twitter
Chenhao Tan, Lillian Lee, and Bo Pang · 2014
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks, 2015
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
CLUTRR: A diagnostic benchmark for inductive reasoning from text
Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Thinking like a skeptic: Defeasible inference in natural language
Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, and Yejin Choi · 2020
Earlier work this paper cites.
The child as hacker: building more human-like models of learning
Joshua Stewart Rule · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Abduction
Igor Douven · 2021
Earlier work this paper cites.
The upworthy research archive, a time series of 32,487 experiments in U.S. media
Jorge Nathan Matias, Kevin Munger, Marianne Aubin Le Quere, and Charles R. Ebersole · 2021
Earlier work this paper cites.
ProofWriter: Generating implications, proofs, and abductive statements over natural language
Oyvind Tafjord, Bhavana Dalvi, and Peter Clark · 2021
Earlier work this paper cites.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi · 2022
Earlier work this paper cites.
Can language models learn from explanations in context?
Andrew Lampinen, Ishita Dasgupta, Stephanie Chan, Kory Mathewson, Mh Tessler, Antonia Creswell, James McClelland, Jane Wang, and Felix Hill · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Reframing human-AI collaboration for generating free-text explanations
Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi · 2022
Earlier work this paper cites.
AbductionRules: Training transformers to explain unexpected inputs
Nathan Young, Qiming Bao, Joshua Bensemann, and Michael Witbrock · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
BIG bench authors · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Self-consistent narrative prompts on abductive natural language inference
Chunkit Chan, Xin Liu, Tsz Ho Chan, Jiayang Cheng, Yangqiu Song, Ginny Wong, and Simon See · 2023
Cited alongside, same era.
True detective: A deep abductive reasoning benchmark undoable for GPT-3 and challenging for GPT-4
Maksym Del and Mark Fishel · 2023
Cited alongside, same era.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang · 2023
Cited alongside, same era.
Brainteaser: Lateral thinking puzzles for large language models, 2023
Yifan Jiang, Filip Ilievski, Kaixin Ma, and Zhivar Sourati · 2023
Can llms follow simple rules?, 2024
Norman Mu, Sarah Chen, Zifan Wang, Sizhe Chen, David Karamardian, Lulwa Aljeraisy, Basel Alomair, Dan Hendrycks, and David Wagner · 2024
Later among the works it cites.
Qian Niu, Junyu Liu, Ziqian Bi, Pohsun Feng, Benji Peng, Keyu Chen, Ming Li, Lawrence KQ Yan, Yichao Zhang, Caitlyn Heqi Yin, Cheng Fei, Tianyang Wang, Yunze Wang, Silin Chen, and Ming Liu · 2024
Later among the works it cites.
TopicGPT: A prompt-based topic modeling framework
Chau Minh Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, and Mohit Iyyer · 2024
Later among the works it cites.
Reasoning with large language models, a survey, 2024
Aske Plaat, Annie Wong, Suzan Verberne, Joost Broekens, Niki van Stein, and Thomas Back · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
StoryAnalogy: Deriving story-level analogies from large language models to unlock analogical understanding
Cheng Jiayang, Lin Qiu, Tsz Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, and Zheng Zhang · 2023
Cited alongside, same era.
Deductive verification of chain-of-thought reasoning
Zhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang, Mingu Lee, Roland Memisevic, and Hao Su · 2023
Cited alongside, same era.
How well do sota legal reasoning models support abductive reasoning?, 2023
Ha-Thanh Nguyen, Randy Goebel, Francesca Toni, Kostas Stathis, and Ken Satoh · 2023
Cited alongside, same era.
Inductive, abductive and deductive theorising
Chitu Okoli · 2023
Cited alongside, same era.
LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers
Theo Olausson, Alex Gu, Ben Lipkin, Cedegao Zhang, Armando Solar-Lezama, Joshua Tenenbaum, and Roger Levy · 2023
Cited alongside, same era.
Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning
Liangming Pan, Alon Albalak, Xinyi Wang, and William Wang · 2023
Cited alongside, same era.
Language models can improve event prediction by few-shot abductive reasoning
Xiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou, James Zhang, Jun Zhou, Chenhao Tan, and Hongyuan Mei · 2023
Cited alongside, same era.
Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, Zhang-Ren Chen, Sihang Zeng, Ermo Hua, Hu Jinfang, and Bowen Zhou · 2024
Later among the works it cites.
Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with hypothesis refinement
Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi, Nouha Dziri, and Xiang Ren · 2024
Later among the works it cites.
Evaluating the deductive competence of large language models
S Seals and Valerie Shalin · 2024
Later among the works it cites.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers, 2024
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto · 2024
Later among the works it cites.
Beyond instruction following: Evaluating inferential rule following of large language models, 2024
Wangtao Sun, Chenxiang Zhang, XueYou Zhang, Xuanqing Yu, Ziyang Huang, Pei Chen, Haotian Xu, Shizhu He, Jun Zhao, and Kang Liu · 2024
Later among the works it cites.
Hypothesis search: Inductive reasoning with language models, 2024
Ruocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu, Nick Haber, and Noah D. Goodman · 2024
Later among the works it cites.
Improving scientific hypothesis generation with knowledge grounded large language models, 2024
Guangzhi Xiong, Eric Xie, Amir Hassan Shariatmadari, Sikun Guo, Stefan Bekiranov, and Aidong Zhang · 2024
Later among the works it cites.
Language models as inductive reasoners
Zonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, and Furu Wei · 2024
Later among the works it cites.
Large language models for automated open-domain scientific hypotheses discovery
Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria · 2024
Later among the works it cites.
UNcommonsense reasoning: Abductive reasoning about uncommon situations
Wenting Zhao, Justin Chiu, Jena Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi, Xiang Li, and Alane Suhr · 2024
Later among the works it cites.
Explaining datasets in words: Statistical models with natural language parameters
Ruiqi Zhong, Heng Wang, Dan Klein, and Jacob Steinhardt · 2024
Later among the works it cites.
Hypothesis generation with large language models
Yangqiaoyu Zhou, Haokun Liu, Tejes Srivastava, Hongyuan Mei, and Chenhao Tan · 2024
Later among the works it cites.
Large language models can learn rules, 2024
Zhaocheng Zhu, Yuan Xue, Xinyun Chen, Denny Zhou, Jian Tang, Dale Schuurmans, and Hanjun Dai · 2024
Later among the works it cites.
A survey on hypothesis generation for scientific discovery in the era of large language models, 2025
Atilla Kaan Alkan, Shashwat Sourav, Maja Jablonska, Simone Astarita, Rishabh Chakrabarty, Nikhil Garuda, Pranav Khetarpal, Maciej Pióro, Dimitrios Tanoglidis, Kartheik G. Iyer, Mugdha S. Polimera, Michael J. Smith, Tirthankar Ghosal, Marc Huertas-Company, Sandor Kruk, Kevin Schawinski, and Ioana Ciucă · 2025
Closest in time.
Agentichypothesis: A survey on hypothesis generation using LLM systems
Adib Bazgir, Rama chandra Praneeth Madugula, and Yuwen Zhang · 2025
Closest in time.
The role of deductive and inductive reasoning in large language models, 2025
Chengkun Cai, Xu Zhao, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, Serge Belongie, and Lei Li · 2025
Closest in time.
Steffen Eger, Yong Cao, Jennifer D’Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, Chenghua Lin, Nafise Sadat Moosavi, Wei Zhao, and Tristan Miller · 2025
Closest in time.
Agentic ai for scientific discovery: A survey of progress, challenges, and future directions, 2025
Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack · 2025
Closest in time.
Inductionbench: Llms fail in the simplest complexity class, 2025
Wenyue Hua, Tyler Wong, Sun Fei, Liangming Pan, Adam Jardine, and William Yang Wang · 2025
Closest in time.
Generating diverse hypotheses for inductive reasoning, 2025
Kang il Lee, Hyukhun Koh, Dongryeol Lee, Seunghyun Yoon, Minsung Kim, and Kyomin Jung · 2025
Closest in time.
Logical reasoning in large language models: A survey, 2025
Hanmeng Liu, Zhizhang Fu, Mengru Ding, Ruoxi Ning, Chaoli Zhang, Xiaozhang Liu, and Yue Zhang · 2025
Closest in time.
Sparse autoencoders for hypothesis generation, 2025
Rajiv Movva, Kenny Peng, Nikhil Garg, Jon Kleinberg, and Emma Pierson · 2025
Closest in time.
Ideasynth: Iterative research idea development through evolving and composing idea facets with literature-grounded feedback
Kevin Pu, KJ Kevin Feng, Tovi Grossman, Tom Hope, Bhavana Dalvi Mishra, Matt Latzke, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue · 2025
Closest in time.
Towards scientific discovery with generative ai: Progress, opportunities, and challenges
Chandan K Reddy and Parshin Shojaee · 2025
Closest in time.
Llm assists hypothesis generation and testing for deliberative questions
Fuchun Wang, Xian Zhou, Wenpeng Hu, Zhunchen Luo, Wei Luo, and Xiaoying Bai · 2025
Closest in time.
Yang Yan, Yu Lu, Renjun Xu, and Zhenzhong Lan · 2025
Closest in time.
Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses, 2025
Zonglin Yang, Wanhao Liu, Ben Gao, Tong Xie, Yuqiang Li, Wanli Ouyang, Soujanya Poria, Erik Cambria, and Dongzhan Zhou · 2025
Closest in time.
Defeasible visual entailment: Benchmark, evaluator, and reward-driven optimization
Yue Zhang, Liqiang Jing, and Vibhav Gogate · 2025
Closest in time.