Fetching the paper…
Reading the bibliography…
Recent work has demonstrated that semantics specified by pretraining data influence how representations of different concepts are organized in a large language model (LLM).
On a theorem of weyl concerning eigenvalues of linear transformations i
Ky Fan · 1949
Earlier work this paper cites.
How to draw a graph
William Thomas Tutte · 1963
Earlier work this paper cites.
Conceptual role semantics
Gilbert Harman · 1982
Earlier work this paper cites.
Semantics, conceptual role
Ned Block · 1998
Earlier work this paper cites.
The Structure and Function of Complex Networks
M. E. J. Newman · 2003
Earlier work this paper cites.
Compression in visual working memory: using statistical regularities to form more efficient memory representations
Timothy F Brady, Talia Konkle, and George A Alvarez · 2009
Earlier work this paper cites.
Percolation on bipartite scale-free networks
H. Hooyberghs, B. Van Schaeybroeck, and J. O. Indekeu · 2010
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
A map of abstract relational knowledge in the human hippocampal–entorhinal cortex
Mona M Garvert, Raymond J Dolan, and Timothy EJ Behrens · 2017
Earlier work this paper cites.
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Spectral and algebraic graph theory
Daniel Spielman · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Transferring structural knowledge across cognitive maps in humans and models
Shirley Mark, Rani Moran, Thomas Parr, Steve W Kennerley, and Timothy EJ Behrens · 2020
Earlier work this paper cites.
The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation
James CR Whittington, Timothy H Muller, Shirley Mark, Guifen Chen, Caswell Barry, Neil Burgess, and Timothy EJ Behrens · 2020
Earlier work this paper cites.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard · 2021
Earlier work this paper cites.
Implicit representations of meaning in neural language models
Belinda Z Li, Maxwell Nye, and Jacob Andreas · 2021
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2022
Earlier work this paper cites.
In-context reinforcement learning with algorithm distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen, Angelos Filos, Ethan Brooks, et al · 2022
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Cited alongside, same era.
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Cited alongside, same era.
Can neural networks learn implicit logic from physical reasoning?
Aaron Traylor, Roman Feiman, and Ellie Pavlick · 2022
Cited alongside, same era.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda · 2024
Closest in time.
Recurrent neural networks learn to store and generate sequences using non-linear representations
Róbert Csordás, Christopher Potts, Christopher D Manning, and Atticus Geiger · 2024
Closest in time.
The llama 3 herd of models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, and et al · 2024
Closest in time.
Not all language model features are linear, 2024
Joshua Engels, Isaac Liao, Eric J. Michaud, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Nnsight and ndif: Democratizing access to foundation model internals, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Transformers from an optimization perspective
Yongyi Yang, David P Wipf, et al · 2022
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models, 2023
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Cited alongside, same era.
In-context learning dynamics with random binary sequences
Eric J Bigelow, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, and Tomer D Ullman · 2023
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes, 2023
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Cited alongside, same era.
In-context learning creates task vectors
Roee Hendel, Mor Geva, and Amir Globerson · 2023
Cited alongside, same era.
Jaden Fiotto-Kaufman, Alexander R Loftus, Eric Todd, Jannik Brinkmann, Caden Juang, Koyena Pal, Can Rager, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Michael Ripa, Adam Belfki, Nikhil Prakash, Sumeet Multani, Carla Brodley, Arjun Guha, Jonathan Bell, Byron Wallace, and David Bau · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Gemma Team · 2024
Closest in time.
Abrupt learning in transformers: A case study on matrix completion
Pulkit Gopalani, Ekdeep Singh Lubana, and Wei Hu · 2024
Closest in time.
Evidence of learned look-ahead in a chess-playing neural network
Erik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen, Scott Emmons, and Stuart Russell · 2024
Closest in time.
Towards an understanding of stepwise inference in transformers: A synthetic graph navigation model
Mikail Khona, Maya Okawa, Jan Hula, Rahul Ramesh, Kento Nishi, Robert Dick, Ekdeep Singh Lubana, and Hidenori Tanaka · 2024
Closest in time.
A percolation model of emergence: Analyzing transformers trained on a formal language
Ekdeep Singh Lubana, Kyogo Kawaguchi, Robert P Dick, and Hidenori Tanaka · 2024
Closest in time.
Flexible neural representations of abstract structural knowledge in the human entorhinal cortex
Shirley Mark, Phillipp Schwartenbeck, Avital Hahamy, Veronika Samborska, Alon B Baram, and Timothy E Behrens · 2024
Closest in time.
Samuel Marks and Max Tegmark · 2024
Closest in time.
Representation shattering in transformers: A synthetic study with knowledge editing
Kento Nishi, Maya Okawa, Rahul Ramesh, Mikail Khona, Ekdeep Singh Lubana, and Hidenori Tanaka · 2024
Closest in time.
Why think step by step? reasoning emerges from the locality of experience
Ben Prystawski, Michael Li, and Noah Goodman · 2024
Closest in time.
Sometimes i am a tree: Data drives unstable hierarchical generalization
Tian Qin, Naomi Saphra, and David Alvarez-Melis · 2024
Closest in time.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner · 2024
Closest in time.
Transformers represent belief state geometry in their residual stream
Adam S Shai, Sarah E Marzen, Lucas Teixeira, Alexander Gietelink Oldenziel, and Paul M Riechers · 2024
Closest in time.
Evaluating the world model implicit in a generative model
Keyon Vafa, Justin Y Chen, Jon Kleinberg, Sendhil Mullainathan, and Ashesh Rambachan · 2024
Closest in time.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski · 2024
Closest in time.