Fetching the paper…
Reading the bibliography…
There have been a huge number of benchmarks proposed to evaluate how large language models (LLMs) behave for logic inference tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
The argument reasoning comprehension task: Identification and reconstruction of implicit warrants
Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, and Benno Stein. 2018 · 1940
Earlier work this paper cites.
Learning syllogism with Euler neural-networks
Tiansi Dong, Chengjiang Li, Christian Bauckhage, Juanzi Li, Stefan Wrobel, and Armin B. Cremers. 2020 · 2007
Earlier work this paper cites.
The Art of Reasoning: An Introduction to Logic and Critical Thinking, Fourth Edition
David Kelley. 2013 · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
A Concise Introduction to Logic
Patrick J. Hurley and Lori Watson. 2018 · 2018
Earlier work this paper cites.
Introduction to Logic
Irving M. Copi, Carl Cohen, and Kenneth D. McMahon. 2019 · 2019
Earlier work this paper cites.
CLUTRR: A diagnostic benchmark for inductive reasoning from text
Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton. 2019 · 2019
Earlier work this paper cites.
LogiQA: A challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020 · 2020
Earlier work this paper cites.
Reclor: A reading comprehension dataset requiring logical reasoning
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020 · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Exploring reasoning schemes: A dataset for syllogism figure identification
Shiya Peng, Lu Liu, Chang Liu, and Dong Yu. 2021 · 2021
Cited alongside, same era.
Avicenna: A challenge dataset for natural language generation toward commonsense syllogistic reasoning
Zeinab Aghahadi and Alireza Talebpour. 2022 · 2022
Cited alongside, same era.
Generalized quantifiers as a source of error in multilingual NLU benchmarks
Ruixiang Cui, Daniel Hershcovich, and Anders Søgaard. 2022 · 2022
Cited alongside, same era.
FOLIO: Natural language reasoning with first-order logic
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, David Peng, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Shafiq Joty, Alexander R. Fabbri, Wojciech Kryscinski, Xi Victoria Lin, Caiming Xiong, and Dragomir Radev. 2022 · 2022
Language models show human-like content effects on reasoning tasks
Andrew K. Lampinen, Ishita Dasgupta, Stephanie C. Y. Chan, Hannah R. Sheahan, Antonia Creswell, Dharshan Kumaran, James L. McClelland, and Felix Hill. 2023 · 2023
Later among the works it cites.
Not all quantifiers are equal: Probing transformer-based language models’ understanding of generalised quantifiers
Tharindu Madusanka, Iqra Zahid, Hao Li, Ian Pratt-Hartmann, and Riza Batista-Navarro. 2023 · 2023
Later among the works it cites.
LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers
Theo Olausson, Alex Gu, Ben Lipkin, Cedegao Zhang, Armando Solar-Lezama, Joshua Tenenbaum, and Roger Levy. 2023 · 2023
Later among the works it cites.
Certified deductive reasoning with language models
Gabriel Poesia, Kanishk Gandhi, Eric Zelikman, and Noah D. Goodman. 2023 · 2023
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Evaluating large language models with NeuBAROCO: Syllogistic reasoning ability and human-like biases
Risako Ando, Takanobu Morishita, Hirohiko Abe, Koji Mineshima, and Mitsuhiro Okada. 2023 · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023 · 2023
Cited alongside, same era.
Legal syllogism prompting: Teaching large language models for legal judgment prediction
Cong Jiang and Xiaolei Yang. 2023 · 2023
Cited alongside, same era.
Hence, socrates is mortal: A benchmark for natural language syllogistic reasoning
Yongkang Wu, Meng Han, Yutao Zhu, Lei Li, Xinyu Zhang, Ruofei Lai, Xiaoguang Li, Yuanhang Ren, Zhicheng Dou, and Zhao Cao. 2023 · 2023
Later among the works it cites.
A systematic comparison of syllogistic reasoning in humans and language models
Tiwalayo Eisape, MH Tessler, Ishita Dasgupta, Fei Sha, Sjoerd van Steenkiste, and Tal Linzen. 2024 · 2024
Closest in time.
GPT-4 technical report
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, et al. 2024 · 2024
Closest in time.
Faithful logical reasoning via symbolic chain-of-thought
Jundong Xu, Hao Fei, Liangming Pan, Qian Liu, Mong-Li Lee, and Wynne Hsu. 2024 · 2024
Closest in time.