Fetching the paper…
Reading the bibliography…
Researchers have recently suggested that models share common representations.
Nonlinear dimensionality reduction by locally linear embedding
Sam T Roweis and Lawrence K Saul · 2000
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Understanding image representations by measuring their equivariance and equivalence
Karel Lenc and Andrea Vedaldi · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning bilingual word embeddings with (almost) no bilingual data
Mikel Artetxe, Gorka Labaka, and Eneko Agirre · 2017
Earlier work this paper cites.
Residual connections encourage iterative inference
Stanislaw Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi · 2017
Earlier work this paper cites.
Gromov-wasserstein alignment of word embedding spaces
David Alvarez-Melis and Tommi Jaakkola · 2018
Earlier work this paper cites.
Adversarial reprogramming of neural networks
Gamaleldin F Elsayed, Ian Goodfellow, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
Factors influencing the surprising instability of word embeddings
Laura Wendlandt, Jonathan K. Kummerfeld, and Rada Mihalcea · 2018
Earlier work this paper cites.
Closed form word embedding alignment
Sunipa Dev, Safia Hassan, and Jeff M Phillips · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Emerging cross-lingual structure in pretrained language models
Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Interpreting gpt: The logit lens, 2020
Nostalgebraist · 2020
Cited alongside, same era.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, XY Han, and David L Donoho · 2020
Cited alongside, same era.
Revisiting model stitching to compare neural representations
Yamini Bansal, Preetum Nakkiran, and Boaz Barak · 2021
Cited alongside, same era.
Analyzing the surprising variability in word embedding stability across languages
Laura Burdick, Jonathan K. Kummerfeld, and Rada Mihalcea · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
The representation landscape of few-shot learning and fine-tuning in large language models
Diego Doimo, Alessandro Serra, Alessio Ansuini, and Alberto Cazzaniga · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Later among the works it cites.
Universal neurons in gpt2 language models
Wes Gurnee, Theo Horsley, Zifan Carl Guo, Tara Rezaei Kheirkhah, Qinyi Sun, Will Hathaway, Neel Nanda, and Dimitris Bertsimas · 2024
Later among the works it cites.
The platonic representation hypothesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contrastive learning inverts the data generating process
Roland S Zimmermann, Yash Sharma, Steffen Schneider, Matthias Bethge, and Wieland Brendel · 2021
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al · 2022
Cited alongside, same era.
Embedding comparator: Visualizing differences in global structure and local neighborhoods via small multiples
Angie Boggust, Brandon Carter, and Arvind Satyanarayan · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Cited alongside, same era.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Cited alongside, same era.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Cited alongside, same era.
Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola · 2024
Later among the works it cites.
Fishing for magikarp: Automatically detecting under-trained tokens in large language models
Sander Land and Max Bartolo · 2024
Later among the works it cites.
A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea · 2024
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2024
Later among the works it cites.
Circuit component reuse across tasks in transformer language models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2024
Later among the works it cites.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
Linguistic collapse: Neural collapse in (large) language models
Robert Wu and Vardan Papyan · 2024
Later among the works it cites.
Implicit geometry of next-token prediction: From language sparsity patterns to model representations
Yize Zhao, Tina Behnia, Vala Vakilian, and Christos Thrampoulidis · 2024
Later among the works it cites.
Algorithmic capabilities of random transformers
Ziqian Zhong and Jacob Andreas · 2024
Later among the works it cites.
Activation space interventions can be transferred between large language models
Narmeen Fatimah Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Michael Lan, Abir HARRASSE, and Amir Abdullah · 2025
Closest in time.