Fetching the paper…
Reading the bibliography…
We propose WorldSense, a benchmark designed to assess the extent to which LLMs are consistently able to sustain tacit world models, by testing how they draw simple inferences from descriptions of simple arrangements of entities.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Transitive inferences and memory in young children
Peter E Bryant and Thomas Trabasso. 1971 · 1971
Earlier work this paper cites.
Reasoning in the chimpanzee: Ii. transitive inference
Douglas J Gillan. 1981 · 1981
Earlier work this paper cites.
Transitive inference formation in pigeons
Lorenzo Von Fersen, Clive D Wynne, Juan D Delius, and John E Staddon. 1991 · 1991
Earlier work this paper cites.
Transitive inference in rats: A test of the spatial coding hypothesis
William A Roberts and Maria T Phelps. 1994 · 1994
Earlier work this paper cites.
Fish can infer social rank by observation alone
Logan Grosenick, Tricia S Clement, and Russell D Fernald. 2007 · 2007
Earlier work this paper cites.
Nonverbal transitive inference: Effects of task and awareness on human performance
Olga F Lazareva and Edward A Wasserman. 2010 · 2010
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, and Tomás Mikolov. 2016 · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Cited alongside, same era.
Eight things to know about large language models
Samuel R Bowman. 2023 · 2023
Closest in time.
Language models represent space and time
Wes Gurnee and Max Tegmark. 2023 · 2023
Closest in time.
The debate over understanding in AI’s large language models
Melanie Mitchell and David C. Krakauer. 2023 · 2023
Closest in time.
Evaluating task understanding through multilingual consistency: A chatgpt case study
Xenia Ohmer, Elia Bruni, and Dieuwke Hupkes. 2023 · 2023
Closest in time.
Gpt4geo: How a language model sees the world’s geography
Jonathan Roberts, Timo Lüddecke, Sowmen Das, Kai Han, and Samuel Albanie. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Machine translation for a very low-resource language - layer freezing approach on transfer learning
Amartya Chowdhury, Deepak K. T., Samudra Vijaya K, and S. R. Mahadeva Prasanna. 2022 · 2022
Cited alongside, same era.
Are prompt-based models clueless?
Pride Kavumba, Ryo Takahashi, and Yusuke Oda. 2022 · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022 · 2022
Cited alongside, same era.
Closest in time.
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023 · 2023
Closest in time.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Closest in time.
A causal framework to quantify the robustness of mathematical reasoning with language models
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schoelkopf, and Mrinmaya Sachan. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Lucas Weber, Elia Bruni, and Dieuwke Hupkes. 2023 · 2023
Closest in time.
Don’t make your llm an evaluation benchmark cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. 2023 · 2023
Closest in time.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2094
Closest in time.