Fetching the paper…
Reading the bibliography…
With the increasing interest in using large language models (LLMs) for planning in natural language, understanding their behaviors becomes an important research question.
A Completeness Theorem in Modal Logic
Saul A. Kripke. 1959 · 1959
Earlier work this paper cites.
Semantical Analysis of Modal Logic I Normal Modal Propositional Calculi
Saul A. Kripke. 1963 · 1963
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff. 1978 · 1978
Earlier work this paper cites.
Mental models: Towards a cognitive science of language, inference, and consciousness
Philip Nicholas Johnson-Laird. 1983 · 1983
Earlier work this paper cites.
Does the autistic child have a “theory of mind” ?
Simon Baron-Cohen, Alan M. Leslie, and Uta Frith. 1985 · 1985
Earlier work this paper cites.
Planning as satisfiability
Henry A Kautz, Bart Selman, et al. 1992 · 1992
Earlier work this paper cites.
The problem of logical form equivalence
Stuart M. Shieber. 1993 · 1993
Earlier work this paper cites.
Hierarchical linear models: Applications and data analysis methods
Stephen W Raudenbush. 2002 · 2002
Earlier work this paper cites.
Fitting Linear Mixed-Effects Models Using lme4
Douglas Bates, Martin Mächler, Ben Bolker, and Steve Walker. 2015 · 2015
Earlier work this paper cites.
Machine Theory of Mind
Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, S. M. Ali Eslami, and Matthew Botvinick. 2018 · 2018
Earlier work this paper cites.
When Does a Reasoner Respond: Nothing Follows?: 41st Annual Meeting of the Cognitive Science Society
Marco Ragni, Hannah Dames, Daniel Brand, and Nicolas Riesterer. 2019 · 2019
Earlier work this paper cites.
Two types of higher-order readings of wh-questions
Yimei Xiang. 2019 · 2019
Earlier work this paper cites.
LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020 · 2020
Earlier work this paper cites.
Transformers as soft reasoners over language
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2021 · 2021
Earlier work this paper cites.
Teaching temporal logics to neural networks
Christopher Hahn, Frederik Schmitt, Jens U Kreber, Markus Norman Rabe, and Bernd Finkbeiner. 2021 · 2021
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022 · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022 · 2022
Earlier work this paper cites.
LogicInference: A new Datasaet for Teaching Logical Inference to seq2seq Models
Santiago Ontanon, Joshua Ainslie, Vaclav Cvicek, and Zachary Fisher. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Abulhair Saparov and He He. 2022 · 2022
Cited alongside, same era.
Natural language to code translation with execution
Freda Shi, Daniel Fried, Marjan Ghazvininejad, Luke Zettlemoyer, and Sida I Wang. 2022 · 2022
Cited alongside, same era.
Entailer: Answering questions with faithful and truthful chains of reasoning
Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2022 · 2022
Cited alongside, same era.
Modern origins of modal logic
AI@Meta. 2024 · 2024
Later among the works it cites.
Perceptions of Linguistic Uncertainty by Language Models and Humans
Catarina G. Belem, Markelle Kelly, Mark Steyvers, Sameer Singh, and Padhraic Smyth. 2024 · 2024
Later among the works it cites.
A systematic comparison of syllogistic reasoning in humans and language models
Tiwalayo Eisape, Michael Tessler, Ishita Dasgupta, Fei Sha, Sjoerd Steenkiste, and Tal Linzen. 2024 · 2024
Later among the works it cites.
Open llm leaderboard v2
Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024 · 2024
Later among the works it cites.
FOLIO: Natural Language Reasoning with First-Order Logic
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, and Dragomir Radev. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roberta Ballarin. 2023 · 2023
Cited alongside, same era.
Systematic testing of three language models reveals low language accuracy, absence of response stability, and a yes-response bias
Vittoria Dentella, Fritz Günther, and Evelina Leivada. 2023 · 2023
Cited alongside, same era.
Demystifying prompts in language models via perplexity estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah Smith, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Prompting is not a substitute for probability measurements in large language models
Jennifer Hu and Roger Levy. 2023 · 2023
Cited alongside, same era.
Mistral 7B
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models
Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai. 2023 · 2023
Cited alongside, same era.
Locally typical sampling
Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Conditional and Modal Reasoning in Large Language Models
Wesley H. Holliday, Matthew Mandelkern, and Cedegao E. Zhang. 2024 · 2024
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Later among the works it cites.
Language models, like humans, show content effects on reasoning tasks
Andrew K Lampinen, Ishita Dasgupta, Stephanie C Y Chan, Hannah R Sheahan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. 2024 · 2024
Later among the works it cites.
Embers of autoregression show how large language models are shaped by the problem they are trained to solve
R. Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D. Hardy, and Thomas L. Griffiths. 2024 · 2024
Later among the works it cites.
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Microsoft. 2024 · 2024
Later among the works it cites.
Kanishka Misra and Najoung Kim. 2024 · 2024
Later among the works it cites.
Learning to reason with llms
OpenAI. 2024 · 2024
Later among the works it cites.
LogicBench: Towards systematic evaluation of logical reasoning ability of large language models
Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo, Santosh Mashetty, Arindam Mitra, and Chitta Baral. 2024 · 2024
Later among the works it cites.
Code Llama: Open Foundation Models for Code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2024 · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024 · 2024
Later among the works it cites.
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
Yuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan, Jen-tse Huang, Pinjia He, Wenxiang Jiao, and Michael Lyu. 2024 · 2024
Later among the works it cites.
Shi Zong and Jimmy Lin. 2024 · 2024
Later among the works it cites.