Fetching the paper…
Reading the bibliography…
The ability to build and reason about models of the world is essential for situated language understanding.
Intuitive physics
Michael McCloskey. 1983 · 1983
Earlier work this paper cites.
Foundations of Language: Brain, Meaning, Grammar, Evolution
Ray Jackendoff. 2002 · 2002
Earlier work this paper cites.
Can language restructure cognition? The case for space
Asifa Majid, Melissa Bowerman, Sotaro Kita, Daniel BM Haun, and Stephen C Levinson. 2004 · 2004
Earlier work this paper cites.
Core knowledge
Elizabeth S. Spelke and Katherine D. Kinzler. 2007 · 2007
Earlier work this paper cites.
Recognizing textual entailment: Rational, evaluation and approaches
Ido Dagan, Bill Dolan, Bernardo Magnini, and Dan Roth. 2010 · 2010
Earlier work this paper cites.
The Number Sense: How the Mind Creates Mathematics
Stanislas Dehaene. 2011 · 2011
Earlier work this paper cites.
Quantitative analysis of culture using millions of digitized books
Jean-Baptiste Michel, Yuan Kui Shen, Aviva Presser Aiden, Adrian Veres, Matthew K. Gray, Google Books Team, Joseph P. Pickett, Dale Hoiberg, Dan Clancy, Peter Norvig, et al. 2011 · 2011
Earlier work this paper cites.
The Winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
Peter W. Battaglia, Jessica B. Hamrick, and Joshua B. Tenenbaum. 2013 · 2013
Earlier work this paper cites.
Theory of mind
Stephanie M. Carlson, Melissa A. Koenig, and Madeline B. Harms. 2013 · 2013
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Towards AI-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, and Tomás Mikolov. 2016 · 2016
Earlier work this paper cites.
Are distributional representations ready for the real world? evaluating word vectors for grounded perceptual meaning
Li Lucy and Jon Gauthier. 2017 · 2017
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber. 2018 · 2018
Earlier work this paper cites.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C. Lipton. 2018 · 2018
Earlier work this paper cites.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Distributional semantics as a source of visual knowledge
Molly Lewis, Martin Zettersten, and Gary Lupyan. 2019 · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2020 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M. Bender and Alexander Koller. 2020 · 2020
Cited alongside, same era.
PIQA: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020 · 2020
Cited alongside, same era.
SyntaxGym: An online platform for targeted evaluation of language models
Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy. 2020 · 2020
Cited alongside, same era.
Hyponli: Exploring the artificial patterns of hypothesis-only bias in natural language inference
Tianyu Liu, Zheng Xin, Baobao Chang, and Zhifang Sui. 2020 · 2020
Cited alongside, same era.
Learning as the unsupervised alignment of conceptual systems
Brett D. Roads and Bradley C. Love. 2020 · 2020
Cited alongside, same era.
Do neural language models overcome reporting bias?
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Later among the works it cites.
Event knowledge in large language models: The gap between the impossible and the unlikely
Carina Kauf, Anna A. Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci. 2023 · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5 technical report
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. 2023 · 2023
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vered Shwartz and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Exploring what is encoded in distributional word vectors: A neurobiologically motivated analysis
Akira Utsumi. 2020 · 2020
Cited alongside, same era.
BLiMP: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020 · 2020
Cited alongside, same era.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard. 2021 · 2021
Cited alongside, same era.
The perception of relations
Alon Hafri and Chaz Firestone. 2021 · 2021
Cited alongside, same era.
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Evaluating statistical language models as pragmatic reasoners
Benjamin Lipkin, Lionel Wong, Gabriel Grand, and Josh Tenenbaum. 2023 · 2023
Later among the works it cites.
Mass editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2023 · 2023
Later among the works it cites.
COMPS: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models
Kanishka Misra, Julia Rayz, and Allyson Ettinger. 2023 · 2023
Later among the works it cites.
MPT-30B: Raising the bar for open-source foundation models
MosaicML. 2023 · 2023
Later among the works it cites.
Large language models fail on trivial alterations to theory-of-mind tasks
Tomer Ullman. 2023 · 2023
Later among the works it cites.
Efficient guided generation for large language models
Brandon T Willard and Rémi Louf. 2023 · 2023
Later among the works it cites.
Lionel Wong, Gabriel Grand, Alexander K. Lew, Noah D. Goodman, Vikash K. Mansinghka, Jacob Andreas, and Joshua B. Tenenbaum. 2023 · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team Gemma, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy. 2024 · 2024
Closest in time.
Auxiliary task demands mask the capabilities of smaller language models
Jennifer Hu and Michael Frank. 2024 · 2024
Closest in time.
Language models align with human judgments on key grammatical constructions
Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova, and Roger Levy. 2024 · 2024
Closest in time.
Log probability scores provide a closer match to human plausibility judgments than prompt-based evaluations
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Evelina Fedorenko, and Jacob Andreas. 2024 · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Closest in time.
Log probabilities are a reliable estimate of semantic plausibility in base and instruction-tuned language models
Carina Kauf, Emmanuele Chersoni, Alessandro Lenci, Evelina Fedorenko, and Anna A Ivanova. 2024 · 2024
Closest in time.
Evaluating large language models in theory of mind tasks
Michal Kosinski. 2024 · 2024
Closest in time.
Can language models handle recursively nested grammatical structures? a case study on comparing models and humans
Andrew Lampinen. 2024 · 2024
Closest in time.
Naive psychology depends on naive physics
Shari Liu, Joseph Outa, and Seda Akbiyik. 2024 · 2024
Closest in time.
Dissociating language and thought in large language models
Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2024 · 2024
Closest in time.
Large-scale evidence for logarithmic effects of word predictability on reading time
Cory Shain, Clara Meister, Tiago Pimentel, Ryan Cotterell, and Roger Levy. 2024 · 2024
Closest in time.
Cognitive representations of social relationships and their developmental origins
Ashley J. Thomas. 2024 · 2024
Closest in time.
From task structures to world models: what do llms know?
Ilker Yildirim and LA Paul. 2024 · 2024
Closest in time.
Shades of zero: Distinguishing impossibility from inconceivability
Jennifer Hu, Felix Sosa, and Tomer D. Ullman. 2025 · 2025
Closest in time.