Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have come closest among all models to date to mastering human language, yet opinions about their linguistic and cognitive capabilities remain split.
Computing Machinery and Intelligence
A. M. Turing · 1950
Earlier work this paper cites.
Syntactic Structures
N. Chomsky · 1957
Earlier work this paper cites.
Phonology in generative grammar
M. Halle · 1962
Earlier work this paper cites.
Eliza—a computer program for the study of natural language communication between man and machine
J. Weizenbaum · 1966
Earlier work this paper cites.
Logic and conversation
H. Grice · 1975
Earlier work this paper cites.
Strategies of Discourse Comprehension
T. A. Van Dijk and W. Kintsch · 1983
Earlier work this paper cites.
Lexical Semantics
D. A. Cruse · 1986
Earlier work this paper cites.
Parallel Distributed Processing
D. E. Rumelhart and J. L. McClelland · 1986
Earlier work this paper cites.
On language and connectionism: Analysis of a parallel distributed processing model of language acquisition
S. Pinker and A. Prince · 1988
Earlier work this paper cites.
Functionalism and the competition model
E. Bates and B. MacWhinney · 1989
Earlier work this paper cites.
A framework for the cooperation of learning algorithms
L. Bottou and P. Gallinari · 1990
Earlier work this paper cites.
Linguistics and cognitive science: problems and mysteries
N. Chomsky · 1991
Earlier work this paper cites.
Arenas of Language Use
H. H. Clark · 1992
Earlier work this paper cites.
Why the child’s theory of mind really is a theory
A. Gopnik and H. M. Wellman · 1992
Earlier work this paper cites.
A human universal: the capacity to learn a language
L. R. Gleitman · 1993
Earlier work this paper cites.
Learning and development in neural networks: the importance of starting small
J. Elman · 1993
Earlier work this paper cites.
The lexical nature of syntactic ambiguity resolution
M. C. MacDonald, N. J. Pearlmutter, and M. S. Seidenberg · 1994
Earlier work this paper cites.
The role of language in intelligence
D. C. Dennett · 1994
Earlier work this paper cites.
Statistical learning by 8-month-old infants
J. Saffran, R. Aslin, and E. Newport · 1996
Earlier work this paper cites.
Using Language
H. H. Clark · 1996
Earlier work this paper cites.
Neural networks for modelling and control
E. Ronco and P. J. Gawthrop · 1997
Earlier work this paper cites.
Presumptive Meanings: The Theory of Generalized Conversational Implicature
S. Levinson · 2000
Earlier work this paper cites.
Foundations of Language: Brain, meaning, grammar, evolution, 2002
R. Jackendoff · 2002
Earlier work this paper cites.
Neural systems underlying British Sign Language and audio-visual English processing in native users
M. MacSweeney, B. Woll, R. Campbell, P. K. McGuire, A. S. David, S. C. R. Williams, J. Suckling, G. A. Calvert, and M. J. Brammer · 2002
Earlier work this paper cites.
The cognitive functions of language
P. Carruthers · 2002
Earlier work this paper cites.
What Does the Frontomedian Cortex Contribute to Language Processing: Coherence or Theory of Mind?
E. C. Ferstl and D. Y. von Cramon · 2002
Earlier work this paper cites.
Early word-learning and conceptual development: Everything had a name, and each name gave birth to a new thought
S. R. Waxman · 2002
Earlier work this paper cites.
Voxel-based lesion-symptom mapping
E. Bates, S. M. Wilson, A. P. Saygin, F. Dick, M. I. Sereno, R. T. Knight, and N. F. Dronkers · 2003
Earlier work this paper cites.
People thinking about thinking people. The role of the temporo-parietal junction in "theory of mind"
R. Saxe and N. Kanwisher · 2003
Earlier work this paper cites.
Language and identity
M. Bucholtz and K. Hall · 2004
Earlier work this paper cites.
Uniquely human social cognition
R. Saxe · 2006
Earlier work this paper cites.
It’s the thought that counts: specific brain regions for one component of theory of mind
R. Saxe and L. J. Powell · 2006
Earlier work this paper cites.
Is syntactic knowledge probabilistic? Experiments with the English dative alternation
J. Bresnan · 2007
Earlier work this paper cites.
Where do you know what you know? The representation of semantic knowledge in the human brain
K. Patterson, P. J. Nestor, and T. T. Rogers · 2007
Earlier work this paper cites.
Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition
D. Jurafsky and J. H. Martin · 2009
Earlier work this paper cites.
Language promotes false-belief understanding: Evidence from learners of a new sign language
J. E. Pyers and A. Senghas · 2009
Earlier work this paper cites.
New method for fMRI investigations of language: defining ROIs functionally in individual subjects
E. Fedorenko, P.-J. Hsieh, A. Nieto-Castañón, S. Whitfield-Gabrieli, and N. Kanwisher · 2010
Earlier work this paper cites.
Distributional memory: A general framework for corpus-based semantics
M. Baroni and A. Lenci · 2010
Earlier work this paper cites.
The multiple-demand (MD) system of the primate brain: mental programs for intelligent behaviour
J. Duncan · 2010
Earlier work this paper cites.
Fluid intelligence loss linked to restricted regions of damage within frontal and parietal cortex
A. Woolgar, A. Parr, R. Cusack, R. Thompson, I. Nimmo-Smith, T. Torralva, M. Roca, N. Antoun, F. Manes, and J. Duncan · 2010
Earlier work this paper cites.
Shared language: overlap and segregation of the neuronal infrastructure for speaking and listening revealed by functional MRI
L. Menenti, S. M. E. Gierhan, K. Segaert, and P. Hagoort · 2011
Earlier work this paper cites.
Functional specificity for high-level linguistic processing in the human brain
E. Fedorenko, M. K. Behr, and N. Kanwisher · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Earlier work this paper cites.
Cortical representation of the constituent structure of sentences
C. Pallier, A.-D. Devauchelle, and S. Dehaene · 2011
Earlier work this paper cites.
Topographic Mapping of a Hierarchy of Temporal Receptive Windows Using a Narrated Story
Y. Lerner, C. J. Honey, L. J. Silbert, and U. Hasson · 2011
Earlier work this paper cites.
Thought beyond language: neural dissociation of algebra and natural language
M. M. Monti, L. M. Parsons, and D. N. Osherson · 2012
Earlier work this paper cites.
Vector space models of word meaning and phrase meaning: A survey
K. Erk · 2012
Earlier work this paper cites.
Colorless green ideas learn furiously: Chomsky and the two cultures of statistical learning
P. Norvig · 2012
Earlier work this paper cites.
A pleasant three days in Philadelphia: Arguments for a pseudopartitive analysis
C. Keenan · 2013
Earlier work this paper cites.
Reporting bias and knowledge acquisition
J. Gordon and B. Van Durme · 2013
Earlier work this paper cites.
Distributional Learning as a Theory of Language Acquisition
A. Clark · 2014
Earlier work this paper cites.
Neuropragmatics
P. Hagoort and S. C. Levinson · 2014
Earlier work this paper cites.
Functional Organization of Social Perception and Cognition in the Superior Temporal Sulcus
B. Deen, K. Koldewyn, N. Kanwisher, and R. Saxe · 2015
Earlier work this paper cites.
Structures, not strings: linguistics as part of the cognitive sciences
M. B. Everaert, M. A. Huybregts, N. Chomsky, R. C. Berwick, and J. J. Bolhuis · 2015
Earlier work this paper cites.
Origins of the brain networks for advanced mathematics in expert mathematicians
M. Amalric and S. Dehaene · 2016
Earlier work this paper cites.
Language and thought are not the same thing: evidence from neuroimaging and neurological patients: Language versus thought
E. Fedorenko and R. A. Varley · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch · 2016
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
T. Linzen, E. Dupoux, and Y. Goldberg · 2016
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
A. Ettinger, A. Elgohary, and P. Resnik · 2016
Earlier work this paper cites.
Neural correlate of the construction of sentence meaning
E. Fedorenko, T. L. Scott, P. Brunner, W. G. Coon, B. Pritchett, G. Schalk, and N. Kanwisher · 2016
Earlier work this paper cites.
Functional neuroanatomy of intuitive physical inference
J. Fischer, J. G. Mikhael, J. B. Tenenbaum, and N. Kanwisher · 2016
Earlier work this paper cites.
Localizing Pain Matrix and Theory of Mind networks with both verbal and non-verbal stimuli
N. Jacoby, E. Bruneau, J. Koster-Hale, and R. Saxe · 2016
Earlier work this paper cites.
A new fun and robust version of an fMRI localizer for the frontotemporal language system
T. L. Scott, J. Gallée, and E. Fedorenko · 2017
Earlier work this paper cites.
Discovering Event Structure in Continuous Narrative Perception and Memory
C. Baldassano, J. Chen, A. Zadbood, J. W. Pillow, U. Hasson, and K. A. Norman · 2017
Earlier work this paper cites.
The Contribution of Grammar, Vocabulary and Theory of Mind in Pragmatic Language Competence in Children with Autistic Spectrum Disorders
C. Andrés-Roqueta and N. Katsos · 2017
Earlier work this paper cites.
Attention is All you Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
K. Gulordava, P. Bojanowski, E. Grave, T. Linzen, and M. Baroni · 2018
Earlier work this paper cites.
Fluid intelligence is supported by the multiple-demand system not the language system
A. Woolgar, J. Duncan, F. Manes, and E. Fedorenko · 2018
Earlier work this paper cites.
Discourse-level comprehension engages medial frontal Theory of Mind brain regions even for expository texts
N. Jacoby and E. Fedorenko · 2018
Earlier work this paper cites.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Earlier work this paper cites.
An amazing four doctoral dissertations
M. Dalrymple and T. H. King · 2019
Earlier work this paper cites.
Explain me this: Creativity, competition, and the partial productivity of constructions
A. E. Goldberg · 2019
Earlier work this paper cites.
The Representation of Semantic Information Across Human Cerebral Cortex During Listening Versus Reading Is Invariant to Stimulus Modality
F. Deniz, A. O. Nunez-Elizalde, A. G. Huth, and J. L. Gallant · 2019
Cited alongside, same era.
Language Mapping in Aphasia
S. M. Wilson, D. K. Eriksson, M. Yen, A. T. Demarco, S. M. Schneck, and J. M. Lucanie · 2019
Cited alongside, same era.
Speech-accompanying gestures are not processed by the language-processing mechanisms
O. Jouravlev, D. Zheng, Z. Balewski, A. L. A. Pongos, Z. Levan, S. Goldin-Meadow, and E. Fedorenko · 2019
Cited alongside, same era.
What can linguistics and deep learning contribute to each other? Response to Pater
T. Linzen · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
Brains and algorithms partially converge in natural language processing
C. Caucheteux and J.-R. King · 2022
Later among the works it cites.
Shared computational principles for language processing in humans and deep language models
A. Goldstein, Z. Zada, E. Buchnik, M. Schain, A. Price, B. Aubrey, S. A. Nastase, A. Feder, D. Emanuel, A. Cohen, A. Jansen, H. Gazula, G. Choe, A. Rao, C. Kim, C. Casto, L. Fanda, W. Doyle, D. Friedman, P. Dugan, L. Melloni, R. Reichart, S. Devore, A. Flinker, L. Hasenfratz, O. Levy, A. Hassidim, M. Brenner, Y. Matias, K. A. Norman, O. Devinsky, and U. Hasson · 2022
Later among the works it cites.
What artificial neural networks can tell us about human language acquisition
A. Warstadt and S. R. Bowman · 2022
Later among the works it cites.
Systematic inequalities in language technology performance across the world’s languages
D. Blasi, A. Anastasopoulos, and G. Neubig · 2022
Later among the works it cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
K. Tirumala, A. Markosyan, L. Zettlemoyer, and A. Aghajanyan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Hewitt and C. D. Manning · 2019
Cited alongside, same era.
The emergence of number and syntax units in LSTM language models
Y. Lakretz, G. Kruszewski, T. Desbordes, D. Hupkes, S. Dehaene, and M. Baroni · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
I. Tenney, D. Das, and E. Pavlick · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
T. McCoy, E. Pavlick, and T. Linzen · 2019
Cited alongside, same era.
Quantity doesn’t buy quality syntax with neural language models
M. van Schijndel, A. Mueller, and T. Linzen · 2019
Cited alongside, same era.
What kind of language is hard to language-model?
S. J. Mielke, R. Cotterell, K. Gorman, B. Roark, and J. Eisner · 2019
Cited alongside, same era.
Language Models as Knowledge Bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller · 2019
Cited alongside, same era.
Large language models still can’t plan (a benchmark for llms on planning and reasoning about change)
K. Valmeekam, A. Olmo, S. Sreedharan, and S. Kambhampati · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Later among the works it cites.
Semantic projection recovers rich human knowledge of multiple object features from word embeddings
G. Grand, I. A. Blank, F. Pereira, and E. Fedorenko · 2022
Later among the works it cites.
Things not written in text: Exploring spatial commonsense from visual signals
X. Liu, D. Yin, Y. Feng, and D. Zhao · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, et al · 2022
Later among the works it cites.
HiStruct+: Improving extractive text summarization with hierarchical structure information
Q. Ruan, M. Ostendorff, and G. Rehm · 2022
Later among the works it cites.
When a sentence does not introduce a discourse entity, transformer-based models still sometimes refer to it
S. Schuster and T. Linzen · 2022
Later among the works it cites.
Neural theory-of-mind? on the limits of social intelligence in large LMs
M. Sap, R. Le Bras, D. Fried, and Y. Choi · 2022
Later among the works it cites.
Exact number concepts are limited to the verbal count range
B. Pitt, E. Gibson, and S. T. Piantadosi · 2022
Later among the works it cites.
Relational Memory-Augmented Language Models
Q. Liu, D. Yogatama, and P. Blunsom · 2022
Later among the works it cites.
Brain-like functional specialization emerges spontaneously in deep neural networks
K. Dobs, J. Martinez, A. J. E. Kell, and N. Kanwisher · 2022
Later among the works it cites.
Coordination among neural modules through a shared global workspace
A. Goyal, A. Didolkar, A. Lamb, K. Badola, N. R. Ke, N. Rahaman, J. Binas, C. Blundell, M. Mozer, and Y. Bengio · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon, et al · 2022
Later among the works it cites.
A survey on evaluation of large language models
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al · 2023
Closest in time.
The foundation model transparency index
R. Bommasani, K. Klyman, S. Longpre, S. Kapoor, N. Maslej, B. Xiong, D. Zhang, and P. Liang · 2023
Closest in time.
Why does surprisal from larger transformer-based language models provide a poorer fit to human reading times?
B.-D. Oh and W. Schuler · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Augmented language models: a survey
G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celikyilmaz, et al · 2023
Closest in time.
Precision fmri reveals that the language-selective network supports both phrase-structure building and lexical access during language production
J. Hu, H. Small, H. Kean, A. Takahashi, L. Zekelman, D. Kleinman, E. Ryan, A. Nieto-Castañón, V. Ferreira, and E. Fedorenko · 2023
Closest in time.
The language network is not engaged in object categorization
Y. Benn, A. A. Ivanova, O. Clark, Z. Mineroff, C. Seikus, J. S. Silva, R. Varley, and E. Fedorenko · 2023
Closest in time.
The human language system, including its inferior frontal component in “Broca’s area,” does not support music perception
X. Chen, J. Affourtit, R. Ryskin, T. I. Regev, S. Norman-Haignere, O. Jouravlev, S. Malik-Moraleda, H. Kean, R. Varley, and E. Fedorenko · 2023
Closest in time.
What are large language models supposed to model?
I. A. Blank · 2023
Closest in time.
Computational language modeling and the promise of in silico experimentation
S. Jain, V. A. Vo, L. Wehbe, and A. G. Huth · 2023
Closest in time.
Openly accessible LLMs can help us to understand human cognition
M. C. Frank · 2023
Closest in time.
Understanding natural language understanding systems. a critical analysis
A. Lenci · 2023
Closest in time.
How much do language models copy from their training data? evaluating linguistic novelty in text generation using RAVEN
R. T. McCoy, P. Smolensky, T. Linzen, J. Gao, and A. Celikyilmaz · 2023
Closest in time.
Mean BERTs make erratic language teachers: the effectiveness of latent bootstrapping in low-resource settings
D. Samuel · 2023
Closest in time.
Findings of the BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora
A. Warstadt, A. Mueller, L. Choshen, E. Wilcox, C. Zhuang, J. Ciro, R. Mosquera, B. Paranjabe, A. Williams, T. Linzen, and R. Cotterell · 2023
Closest in time.
COMPS: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models
K. Misra, J. Rayz, and A. Ettinger · 2023
Closest in time.
Construction grammar provides unique insight into neural language models
L. Weissweiler, T. He, N. Otani, D. R. Mortensen, L. Levin, and H. Schütze · 2023
Closest in time.
A discerning several thousand judgments: GPT-3 rates the article + adjective + numeral + noun construction
K. Mahowald · 2023
Closest in time.
Characterizing English Preposing in PP constructions
C. Potts · 2023
Closest in time.
The quantization model of neural scaling
E. J. Michaud, Z. Liu, U. Girit, and M. Tegmark · 2023
Closest in time.
Modern language models refute Chomsky’s approach to language
S. T. Piantadosi · 2023
Closest in time.
R. T. McCoy, S. Yao, D. Friedman, M. Hardy, and T. L. Griffiths · 2023
Closest in time.
How poor is the stimulus? evaluating hierarchical generalization in neural networks trained on child-directed speech
A. Yedetore, T. Linzen, R. Frank, and R. T. McCoy · 2023
Closest in time.
Not all layers are equally as important: Every layer counts BERT
L. Georges Gabriel Charpentier and D. Samuel · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, S. Welleck, P. West, C. Bhagavatula, R. L. Bras, J. D. Hwang, S. Sanyal, X. Ren, A. Ettinger, Z. Harchaoui, and Y. Choi · 2023
Closest in time.
Z. Wu, L. Qiu, A. Ross, E. Akyürek, B. Chen, B. Wang, N. Kim, J. Andreas, and Y. Kim · 2023
Closest in time.
On the paradox of learning to reason from data
H. Zhang, L. H. Li, T. Meng, K.-W. Chang, and G. Van den Broeck · 2023
Closest in time.
L. Wong, G. Grand, A. K. Lew, N. D. Goodman, V. K. Mansinghka, J. Andreas, and J. B. Tenenbaum · 2023
Closest in time.
From task structures to world models: What do LLMs know?
I. Yildirim and L. Paul · 2023
Closest in time.
Evaluating verifiability in generative search engines
N. Liu, T. Zhang, and P. Liang · 2023
Closest in time.
M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr · 2023
Closest in time.
Carpe diem: On the evaluation of world knowledge in lifelong language models
Y. Kim, J. Yoon, S. Ye, S. J. Hwang, and S.-Y. Yun · 2023
Closest in time.
Crawling the internal knowledge-base of language models
R. Cohen, M. Geva, J. Berant, and A. Globerson · 2023
Closest in time.
Entity tracking in language models
N. Kim and S. Schuster · 2023
Closest in time.
Non-literal language processing is jointly supported by the language and theory of mind networks: Evidence from a novel meta-analytic fmri approach
M. Hauptman, I. Blank, and E. Fedorenko · 2023
Closest in time.
A fine-grained comparison of pragmatic language understanding in humans and language models
J. Hu, S. Floyd, O. Jouravlev, E. Fedorenko, and E. Gibson · 2023
Closest in time.
Theory of mind may have spontaneously emerged in large language models
M. Kosinski · 2023
Closest in time.
Large language models fail on trivial alterations to theory-of-mind tasks
T. Ullman · 2023
Closest in time.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
N. Shapira, M. Levy, S. H. Alavi, X. Zhou, Y. Choi, Y. Goldberg, M. Sap, and V. Shwartz · 2023
Closest in time.
Do large language models know what humans know?
S. Trott, C. Jones, T. Chang, J. Michaelov, and B. Bergen · 2023
Closest in time.
Understanding social reasoning in language models with language models
K. Gandhi, J.-P. Fränken, T. Gerstenberg, and N. Goodman · 2023
Closest in time.
Minding language models’ (lack of) theory of mind: A plug-and-play multi-character belief tracker
M. Sclar, S. Kumar, P. West, A. Suhr, Y. Choi, and Y. Tsvetkov · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
LLM+P: Empowering large language models with optimal planning proficiency
B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone · 2023
Closest in time.
Transmission versus truth, imitation versus innovation: What children can do that large language and language-and-vision models cannot (yet)
E. Yiu, E. Kosoy, and A. Gopnik · 2023
Closest in time.
The debate over understanding in AI’s large language models
M. Mitchell and D. C. Krakauer · 2023
Closest in time.
Symbols and grounding in large language models
E. Pavlick · 2023
Closest in time.
D. C. Mollo and R. Millière · 2023
Closest in time.
Noam Chomsky: Language, Cognition, and Deep Learning: Lex Fridman Podcast #53
L. Fridman · 2024
Closest in time.
Artificial neural network language models predict human brain responses to language even after a developmentally realistic amount of training
E. A. Hosseini, M. Schrimpf, Y. Zhang, S. R. Bowman, N. Zaslavsky, and E. Fedorenko · 2024
Closest in time.
Wolfram plugin for chatgpt
Wolfram · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.
H. Lederman and K. Mahowald · 2024
Closest in time.