Fetching the paper…
Reading the bibliography…
As language models (LMs) become integral to fields like healthcare, law, and journalism, their ability to differentiate between fact, belief, and knowledge is essential for reliable decision-making.
Is justified true belief knowledge?
Edmund L Gettier · 1963
Earlier work this paper cites.
Does knowledge entail belief?
David M Armstrong · 1969
Earlier work this paper cites.
Belief, truth and knowledge
David Malet Armstrong · 1973
Earlier work this paper cites.
Does the autistic child have a “theory of mind”?
Simon Baron-Cohen, Alan M Leslie, and Uta Frith · 1985
Earlier work this paper cites.
Young children understand that looking leads to knowing (so long as they are looking into a single barrel)
Chris Pratt and Peter Bryant · 1990
Earlier work this paper cites.
Theory of mind and the development of social-linguistic intelligence in early childhood
Marilyn Shatz · 1994
Earlier work this paper cites.
The ‘seeing-leads-to-knowing’deficit in autism: The pratt and bryant probe
Simon Baron-Cohen and Frances Goodhart · 1994
Earlier work this paper cites.
Imitation as a mechanism of social cognition: Origins of empathy, theory of mind, and the representation of action
Andrew N Meltzoff · 2002
Earlier work this paper cites.
The myth of factive verbs
Allan Hazlett · 2010
Earlier work this paper cites.
Bayesian theory of mind: Modeling joint belief-desire attribution
Chris Baker, Rebecca Saxe, and Joshua Tenenbaum · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Theory of mind
Stephanie M Carlson, Melissa A Koenig, and Madeline B Harms · 2013
Earlier work this paper cites.
Strong epistemic modality in parliamentary discourse
Milica Vukovic · 2014
Earlier work this paper cites.
The Analysis of Knowledge
Jonathan Jenkins Ichikawa and Matthias Steup · 2018
Earlier work this paper cites.
Revisiting the evaluation of theory of mind through question answering
Matthew Le, Y-Lan Boureau, and Maximilian Nickel · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Epistemic Closure
Steven Luper · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Reclor: A reading comprehension dataset requiring logical reasoning
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng · 2020
Earlier work this paper cites.
The epistemologies of breaking news
Mats Ekström, Amanda Ramsälv, and Oscar Westlund · 2021
Earlier work this paper cites.
Knowledge before belief
Jonathan Phillips, Wesley Buckwalter, Fiery Cushman, Ori Friedman, Alia Martin, John Turri, Laurie Santos, and Joshua Knobe · 2021
Earlier work this paper cites.
Logiqa: a challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang · 2021
Earlier work this paper cites.
Detecting changes in help seeker conversations on a suicide prevention helpline during the covid- 19 pandemic: in-depth analysis using encoder representations from transformers
Salim Salmi, Saskia Mérelle, Renske Gilissen, Rob van der Mei, and Sandjai Bhulai · 2022
Earlier work this paper cites.
Neural theory-of-mind? on the limits of social intelligence in large LMs
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi · 2022
Earlier work this paper cites.
On the acquisition of attitude verbs
Valentine Hacquard and Jeffrey Lidz · 2022
Earlier work this paper cites.
Large language models still can’t plan (a benchmark for llms on planning and reasoning about change)
Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2022
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Antonia Creswell, Murray Shanahan, and Irina Higgins · 2022
Earlier work this paper cites.
The curious case of commonsense intelligence
Yejin Choi · 2022
Earlier work this paper cites.
Commonsenseqa 2.0: Exposing the limits of ai through gamification
Alon Talmor, Ori Yoran, Ronan Le Bras, Chandra Bhagavatula, Yoav Goldberg, Yejin Choi, and Jonathan Berant · 2022
Earlier work this paper cites.
Commonsense knowledge reasoning and generation with pre-trained language models: A survey
Prajjwal Bhargava and Vincent Ng · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou · 2022
Earlier work this paper cites.
Do prompt-based models really understand the meaning of their prompts?
Albert Webson and Ellie Pavlick · 2022
Cited alongside, same era.
Foundations of theory of mind and its development in early childhood
Hannes Rakoczy · 2022
Cited alongside, same era.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Cited alongside, same era.
People are eagerly consulting generative ai chatgpt for mental health advice, stressing out ai ethics and ai law
Lance Eliot · 2023
Cited alongside, same era.
Large language models for therapy recommendations across 3 clinical specialties: comparative study
Theresa Isabelle Wilhelm, Jonas Roos, and Robert Kaczmarczyk · 2023
Cited alongside, same era.
AI and journalism: What’s next?, 2023
How Not to Test GPT-3
Gary Marcus and Ernest Davis · 2023
Later among the works it cites.
Towards a holistic landscape of situated theory of mind in large language models
Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai · 2023
Later among the works it cites.
Epistemic language in news headlines shapes readers’ perceptions of objectivity
Aaron Chuey, Yiwei Luo, and Ellen M Markman · 2024
Closest in time.
The diagnostic and triage accuracy of the gpt-3 artificial intelligence model: an observational study
David M Levine, Rudraksh Tuwani, Benjamin Kompa, Amita Varma, Samuel G Finlayson, Ateev Mehrotra, and Andrew Beam · 2024
Closest in time.
Hallucination-free? assessing the reliability of leading ai legal research tools
Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D Manning, and Daniel E Ho · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Caswell · 2023
Cited alongside, same era.
Robot reporters? here’s how news organisations are using ai in journalism
Luke Hurst · 2023
Cited alongside, same era.
Mirror and express owner publishes first articles written using ai
M Sweney · 2023
Cited alongside, same era.
Chatgpt has entered the classroom: how llms could transform education
Andy Extance · 2023
Cited alongside, same era.
M-powering teachers: Natural language processing powered feedback improves 1: 1 instruction and student outcomes
Dorottya Demszky and Jing Liu · 2023
Cited alongside, same era.
FinGPT: Democratizing internet-scale data for financial large language models
Xiao-Yang Liu, Guoxuan Wang, Hongyang Yang, and Daochen Zha · 2023
Cited alongside, same era.
Can chatgpt forecast stock price movements? return predictability and large language models
Alejandro Lopez-Lira and Yuehua Tang · 2023
Cited alongside, same era.
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho · 2024
Closest in time.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al · 2024
Closest in time.
We Asked Every Am Law 100 Law Firm How They’re Using Gen AI. Here’s What We Learned
Justin Henry · 2024
Closest in time.
AI Survey: Where Artificial Intelligence Stands in the Legal Industry, April 2024
Jack Collens, Rachel Reimer, Gerald Schifman, and Pamela Wilkinson · 2024
Closest in time.
Tutor copilot: A human-ai approach for scaling real-time expertise
Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson, Susanna Loeb, and Dora Demszky · 2024
Closest in time.
Researchers built an ’AI Scientist’ – what can it do?
Davide Castelvecchi · 2024
Closest in time.
Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang · 2024
Closest in time.
The ai scientist: Towards fully automated open-ended scientific discovery
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha · 2024
Closest in time.
Ai wisdom: What happens when you ask an algorithm for relationship advice
David Robson · 2024
Closest in time.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz · 2024
Closest in time.
Dissociating language and thought in large language models
Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko · 2024
Closest in time.
How far are large language models from agents with theory-of-mind?, 2024
Pei Zhou, Aman Madaan, Srividya Pranavi Potharaju, Aditya Gupta, Kevin R. McKee, Ari Holtzman, Jay Pujara, Xiang Ren, Swaroop Mishra, Aida Nematzadeh, Shyam Upadhyay, and Manaal Faruqui · 2024
Closest in time.
OpenToM: A comprehensive benchmark for evaluating theory-of-mind reasoning capabilities of large language models
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, and Yulan He · 2024
Closest in time.
Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim · 2024
Closest in time.
Victoria Basmov, Yoav Goldberg, and Reut Tsarfaty · 2024
Closest in time.
Conditional and modal reasoning in large language models
Wesley H Holliday and Matthew Mandelkern · 2024
Closest in time.
Hello GPT-4o, 2024
OpenAI · 2024
Closest in time.
The Claude 3 Model Family: Opus, Sonnet, Haiku, March 2024
Anthropic · 2024
Closest in time.
Llama 3 Model Card, 2024
AI@Meta · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
string2string: A modern python library for string-to-string algorithms
Mirac Suzgun, Stuart Shieber, and Dan Jurafsky · 2024
Closest in time.
Towards understanding factual knowledge of large language models
Xuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo, Lijie Wen, Philip S. Yu, and Zhijiang Guo · 2024
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
Meta-prompting: Enhancing language models with task-agnostic scaffolding
Mirac Suzgun and Adam Tauman Kalai · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Closest in time.
Perceptions to beliefs: Exploring precursory inferences for theory of mind in large language models
Chani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim, Yeon Seonwoo, Yejin Choi, Alice Oh, and Hyunwoo Kim · 2024
Closest in time.
What does that mean? complementizers and epistemic authority
Rebecca Tollan and Bilge Palaz · 2024
Closest in time.