Fetching the paper…
Reading the bibliography…
Large-scale pretrained language models (LMs) are said to ``lack the ability to connect utterances to the world'' (Bender and Koller, 2020), because they do not have ``mental models of the world' '(Mitchell and Krakauer, 2023).
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
A generalized solution of the orthogonal procrustes problem
Peter H Schönemann. 1966 · 1966
Earlier work this paper cites.
Minds, brains, and programs
John R. Searle. 1980 · 1980
Earlier work this paper cites.
Lexical Competence
D. Marconi. 1997 · 1997
Earlier work this paper cites.
Stepping back inside leibniz’s mill
Paul Lodge and Marc Bobro. 1998 · 1998
Earlier work this paper cites.
Holism, conceptual-role semantics, and syntactic semantics
William J. Rapaport. 2002 · 2002
Earlier work this paper cites.
Intuition pumps and the proper use of thought experiments
Elke Brendel. 2004 · 2004
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Learning bilingual lexicons using the visual similarity of labeled web images
Shane Bergsma and Benjamin Van Durme. 2011 · 2011
Earlier work this paper cites.
Babelnet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network
Roberto Navigli and Simone Paolo Ponzetto. 2012 · 2012
Earlier work this paper cites.
Learning image embeddings using convolutional neural networks for improved multi-modal semantics
Douwe Kiela and Léon Bottou. 2014 · 2014
Earlier work this paper cites.
Is this a wampimuk? cross-modal mapping between distributional semantics and the visual world
Angeliki Lazaridou, Elia Bruni, and Marco Baroni. 2014 · 2014
Earlier work this paper cites.
Visual bilingual lexicon induction with transferred ConvNet features
Douwe Kiela, Ivan Vulić, and Stephen Clark. 2015 · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Multi-modal representations for improved bilingual lexicon learning
Ivan Vulić, Douwe Kiela, Stephen Clark, and Marie-Francine Moens. 2016 · 2016
Earlier work this paper cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018 · 2018
Earlier work this paper cites.
Word Translation Without Parallel Data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Earlier work this paper cites.
Why is unsupervised alignment of English embeddings from different algorithms so hard?
Mareike Hartmann, Yova Kementchedjhieva, and Anders Søgaard. 2018 · 2018
Earlier work this paper cites.
Limitations of cross-lingual learning from image search
Mareike Hartmann and Anders Søgaard. 2018 · 2018
Earlier work this paper cites.
An iterative closest point method for unsupervised word translation
Yedid Hoshen and Lior Wolf. 2018 · 2018
Earlier work this paper cites.
NORMA: Neighborhood sensitive maps for multilingual word embeddings
Ndapa Nakashole. 2018 · 2018
Cited alongside, same era.
Brain-score: Which artificial neural network for object recognition is most brain-like?
Martin Schrimpf, Jonas Kubilius, Ha Hong, Najib Majaj, Rishi Rajalingham, Elias B. Issa, Kohitij Kar, Pouya Bashivan, Jonathan Prescott-Roy, Kailyn Schmidt, Daniel L. K. Yamins, and James J. DiCarlo. 2018 · 2018
Cited alongside, same era.
Representation in Cognitive Science
Nicholas Shea. 2018 · 2018
Cited alongside, same era.
On the Limitations of Unsupervised Bilingual Dictionary Induction
Anders Søgaard, Sebastian Ruder, and Ivan Vulić. 2018 · 2018
Cited alongside, same era.
Predictive processing and the representation wars
Daniel Williams. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Analogy training multilingual encoders
Nicolas Garneau, Mareike Hartmann, Anders Sandholm, Sebastian Ruder, Ivan Vulić, and Anders Søgaard. 2021 · 2021
Later among the works it cites.
Thinking ahead: spontaneous prediction in context as a keystone of language in humans and machines
Ariel Goldstein, Zaid Zada, Eliav Buchnik, Mariano Schain, Amy Price, Bobbi Aubrey, Samuel A. Nastase, Amir Feder, Dotan Emanuel, Alon Cohen, Aren Jansen, Harshvardhan Gazula, Gina Choe, Aditi Rao, Se Catherine Kim, Colton Casto, Lora Fanda, Werner Doyle, Daniel Friedman, Patricia Dugan, Lucia Melloni, Roi Reichart, Sasha Devore, Adeen Flinker, Liat Hasenfratz, Omer Levy, Avinatan Hassidim, Michael Brenner, Yossi Matias, Kenneth A. Norman, Orrin Devinsky, and Uri Hasson. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Later among the works it cites.
The singleton fallacy: Why current critiques of language models miss the point
Magnus Sahlgren and Fredrik Carlsson. 2021 · 2021
Later among the works it cites.
The neural architecture of language: Integrative modeling converges on predictive processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
From brain space to distributional space: The perilous journeys of fMRI decoding
Gosse Minnema and Aurélie Herbelot. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)
Mariya Toneva and Leila Wehbe. 2019 · 2019
Cited alongside, same era.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M. Bender and Alexander Koller. 2020 · 2020
Cited alongside, same era.
Non-linear instance-based cross-lingual mapping for non-isomorphic embedding spaces
Goran Glavaš and Ivan Vulić. 2020 · 2020
Cited alongside, same era.
Martin Schrimpf, Idan Blank, Greta Tuckute, Carina Kauf, Eghbal A. Hosseini, Nancy Kanwisher, Joshua Tenenbaum, and Evelina Fedorenko. 2021 · 2021
Later among the works it cites.
Segformer: Simple and efficient design for semantic segmentation with transformers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. 2021 · 2021
Later among the works it cites.
Predictive Coding or Just Feature Discovery? An Alternative Account of Why Language Models Fit Brain Data
Richard Antonello and Alexander Huth. 2022 · 2022
Later among the works it cites.
Long-range and hierarchical language predictions in brains and algorithms
Charlotte Caucheteux, Alexandre Gramfort, and Jean-Rémi King. 2022 · 2022
Later among the works it cites.
Brains and algorithms partially converge in natural language processing
Charlotte Caucheteux and Jean-Rémi King. 2022 · 2022
Later among the works it cites.
The combination of hebbian and predictive plasticity learns invariant object representations in deep sensory networks
Manu Srinath Halvagal and Friedemann Zenke. 2022 · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022 · 2022
Later among the works it cites.
Meaning without reference in large language models
Steven Piantadosi and Felix Hill. 2022 · 2022
Later among the works it cites.
Emergent structures and training dynamics in large language models
Ryan Teehan, Miruna Clinciu, Oleg Serikov, Eliza Szczechla, Natasha Seelam, Shachar Mirkin, and Aaron Gokaslan. 2022 · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022 · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Structural similarities between language models and neural response measurements
Jiaang Li, Antonia Karamolegkou, Yova Kementchedjhieva, Mostafa Abdou, Sune Lehmann, and Anders Søgaard. 2023 · 2023
Closest in time.
Matthew Mandelkern and Tal Linzen. 2023 · 2023
Closest in time.
A sentence is worth a thousand pictures: Can large language models understand human language?
Gary Marcus, Evelina Leivada, and Elliot Murphy. 2023 · 2023
Closest in time.
Linearly mapping from image to text space
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick. 2023 · 2023
Closest in time.
The debate over understanding in ai’s large language models
Melanie Mitchell and David C. Krakauer. 2023 · 2023
Closest in time.
Dimitri Coelho Mollo and Raphaël Millière. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
The platonic representation hypothesis
Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. 2024 · 2024
Closest in time.