Fetching the paper…
Reading the bibliography…
Recent work found high mutual information between the learned representations of large language models (LLMs) and the geospatial property of its input, hinting an emergent internal model of space.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Cross-validatory choice and assessment of statistical predictions
M. Stone. 1974 · 1974
Earlier work this paper cites.
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et al. 2020 · 2004
Earlier work this paper cites.
Analyzing microarray gene expression data
Geoffrey J. McLachlan, Kim-Anh Do, and Christophe Ambroise. 2004 · 2004
Earlier work this paper cites.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2004
Earlier work this paper cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard H. Hovy. 2020 · 2005
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2006
Earlier work this paper cites.
Learning deep architectures for ai
Yoshua Bengio. 2007 · 2007
Earlier work this paper cites.
Understanding spatial relations through multiple modalities
Soham Dan, Hangfeng He, and Dan Roth. 2020 · 2007
Earlier work this paper cites.
Representational similarity analysis - connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini. 2008 · 2008
Earlier work this paper cites.
A Non-singular Horizontal Position Representation
Kenneth Gade. 2010 · 2010
Cited alongside, same era.
F.w. bessel (1825): The calculation of longitude and latitude from geodesic measurements
C.F.F. Karney and R.E. Deakin. 2010 · 2010
Cited alongside, same era.
Unsupervised feature learning and deep learning: A review and new perspectives
Yoshua Bengio, Aaron C. Courville, and Pascal Vincent. 2012 · 2012
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
The power of deeper networks for expressing natural functions
David Rolnick and Max Tegmark. 2017 · 2017
Cited alongside, same era.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021 · 2021
Later among the works it cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Causalm: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2021 · 2021
Later among the works it cites.
Implicit representations of meaning in neural language models
Belinda Z Li, Maxwell Nye, and Jacob Andreas. 2021 · 2021
Later among the works it cites.
Provable limitations of acquiring meaning from ungrounded form: What will future language models understand?
William Merrill, Yoav Goldberg, Roy Schwartz, and Noah A Smith. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Mario Giulianelli, Jacqueline Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Climbing towards nlu: On meaning, form, and understanding in the age of data
Emily M Bender and Alexander Koller. 2020 · 2020
Cited alongside, same era.
Bert-based spatial information extraction
Hyeong Jin Shin, Jeong Yeon Park, Dae Bum Yuk, and Jae Sung Lee. 2020 · 2020
Cited alongside, same era.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard. 2021 · 2021
Cited alongside, same era.
Mycal Tucker, Peng Qian, and Roger Levy. 2021 · 2021
Later among the works it cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun. 2022 · 2022
Later among the works it cites.
Language models represent space and time
Wes Gurnee and Max Tegmark. 2023 · 2023
Closest in time.
Linearity of relation decoding in transformer language models
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.