Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved remarkable performance on a variety of natural language understanding tasks.
Ro{bert}a: A robustly optimized {bert} pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 1907
Earlier work this paper cites.
A formula for predicting readability
Edgar Dale and Jeanne S. Chall. 1948 · 1948
Earlier work this paper cites.
Two sets of perfect syllogisms
Anne Lehman. 1973 · 1973
Earlier work this paper cites.
Events in the Semantics of English
Terence Parsons. 1990 · 1990
Earlier work this paper cites.
Readability Revisited: The New Dale-Chall Readability Formula
Edgar Dale and Jeanne S. Chall. 1995 · 1995
Earlier work this paper cites.
Using the framework
Robin Cooper, Dick Crouch, Jan Van Eijck, Chris Fox, Johan Van Genabith, Jan Jaspars, Hans Kamp, David Milward, Manfred Pinkal, Massimo Poesio, et al. 1996 · 1996
Earlier work this paper cites.
105The Logical Form of Action Sentences
Donald Davidson. 2001 · 2001
Earlier work this paper cites.
Transformers as soft reasoners over language
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020 · 2002
Earlier work this paper cites.
Isabelle/Hol a Proof Assistant for Higher-Order Logic
Tobias Nipkow, Lawrence C. Paulson, and Markus Wenzel. 2002 · 2002
Earlier work this paper cites.
Exploring neural models for parsing natural language into first-order logic
Hrituraj Singh, Milan Aggrawal, and Balaji Krishnamurthy. 2020 · 2002
Earlier work this paper cites.
Generics: Cognition and Acquisition
Sarah-Jane Leslie. 2008 · 2008
Earlier work this paper cites.
Prover9 and mace4
W. McCune. 2005–2010 · 2010
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach , 3 edition
Stuart Russell and Peter Norvig. 2010 · 2010
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
The TPTP Problem Library and Associated Infrastructure. From CNF to TH0, TPTP v6.4.0
G. Sutcliffe. 2017 · 2017
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
CLUTRR: A diagnostic benchmark for inductive reasoning from text
Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton. 2019 · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
HybridQA: A dataset of multi-hop question answering over tabular and textual data
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. 2020 · 2020
Cited alongside, same era.
Leap-of-thought: Teaching pre-trained models to systematically reason over implicit knowledge
Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, and Jonathan Berant. 2020 · 2020
Cited alongside, same era.
Reclor: A reading comprehension dataset requiring logical reasoning
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020 · 2020
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Closest in time.
On the advance of making language models better reasoners
Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022 · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022 · 2022
Closest in time.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022 · 2022
Closest in time.
Evaluating the rationale understanding of critical reasoning in logical reading comprehension
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformers as soft reasoners over language
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2021 · 2021
Cited alongside, same era.
WinoLogic: A zero-shot logic-based diagnostic dataset for Winograd Schema Challenge
Weinan He, Canming Huang, Yongmei Liu, and Xiaodan Zhu. 2021 · 2021
Cited alongside, same era.
Logiqa: a challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2021 · 2021
Cited alongside, same era.
RuleBERT: Teaching soft rules to pre-trained language models
Mohammed Saeed, Naser Ahmadi, Preslav Nakov, and Paolo Papotti. 2021 · 2021
Cited alongside, same era.
ProofWriter: Generating implications, proofs, and abductive statements over natural language
Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2021 · 2021
Cited alongside, same era.
Diagnosing the first-order logical reasoning ability through LogicNLI
Jidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao, Hao He, and Yaohui Jin. 2021 · 2021
Cited alongside, same era.
Linguistic complexity loss in text-based therapy
Jason Wei, Kelly Finn, Emma Templeton, Thalia Wheatley, and Soroush Vosoughi. 2021 · 2021
Cited alongside, same era.
Akira Kawabata and Saku Sugawara. 2023 · 2023
Closest in time.
Boardgameqa: A dataset for natural language reasoning with contradictory information
Mehran Kazemi, Quan Yuan, Deepti Bhatia, Najoung Kim, Xin Xu, Vaiva Imbrasaite, and Deepak Ramachandran. 2023 · 2023
Closest in time.
L2ceval: Evaluating language-to-code generation capabilities of large language models
Ansong Ni, Pengcheng Yin, Yilun Zhao, Martin Riddell, Troy Feng, Rui Shen, Stephen Yin, Ye Liu, Semih Yavuz, Caiming Xiong, Shafiq Joty, Yingbo Zhou, Dragomir Radev, and Arman Cohan. 2023 · 2023
Closest in time.
LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers
Theo Olausson, Alex Gu, Ben Lipkin, Cedegao Zhang, Armando Solar-Lezama, Joshua Tenenbaum, and Roger Levy. 2023 · 2023
Closest in time.
OpenAI, Josh Achiam, and Others. 2023 · 2023
Closest in time.
Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning
Liangming Pan, Alon Albalak, Xinyi Wang, and William Wang. 2023 · 2023
Closest in time.
Language models can (kind of) reason: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He. 2023 · 2023
Closest in time.
Misery loves complexity: Exploring linguistic complexity in the context of emotion detection
Pranaydeep Singh, Luna De Bruyne, Orphée De Clercq, and Els Lefever. 2023 · 2023
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, and +447 Authors. 2023 · 2023
Closest in time.
Hongda Sun, Weikai Xu, Wei Liu, Jian Luan, Bin Wang, Shuo Shang, Ji-Rong Wen, and Rui Yan. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R Narasimhan. 2023 · 2023
Closest in time.