Fetching the paper…
Reading the bibliography…
We employ new tools from mechanistic interpretability in order to ask whether the internal structure of large language models (LLMs) shows correspondence to the linguistic structures which underlie the languages on which they are trained.
Universal grammar, statistics or both?
Charles D Yang · 2004
Earlier work this paper cites.
Statistics (international student edition)
David Freedman, Robert Pisani, and Roger Purves · 2007
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
A Conneau · 2019
Earlier work this paper cites.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Beto, bentz, becas: The surprising cross-lingual effectiveness of bert
Shijie Wu and Mark Dredze · 2019
Earlier work this paper cites.
Interpreting gpt: the logit lens, 2020
Nostalgebraist · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
Cpm: A large-scale generativ e chinese pre-tned language model
Zhengyan Zhang, Xu Han, Hao Zhou, Ke Pei, Yuxian Gu, Deming Ye, Yujia Qin, Yusheng Su, Haozhe Ji, Jian Guan, et al · 2020
Earlier work this paper cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Earlier work this paper cites.
Mechanisms for handling nested dependencies in neural-network language models and humans
Yair Lakretz, Dieuwke Hupkes, Alessandra Vergallito, Marco Marelli, Marco Baroni, and Stanislas Dehaene · 2021
Earlier work this paper cites.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldwosky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Earlier work this paper cites.
A balanced data approach for evaluating cross-lingual transfer: Mapping the linguistic blood bank
Dan Malkin, Tomasz Limisiewicz, and Gabriel Stanovsky · 2022
Earlier work this paper cites.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Earlier work this paper cites.
Neuron-level interpretation of deep nlp models: A survey, 2022
Hassan Sajjad, Nadir Durrani, and Fahim Dalvi · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Earlier work this paper cites.
Cross-lingual few-shot learning on unseen languages
Genta Winata, Shijie Wu, Mayank Kulkarni, Thamar Solorio, and Daniel Preoţiuc-Pietro · 2022
Earlier work this paper cites.
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al · 2022
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens, 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Pythia: a suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal · 2023
Cited alongside, same era.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Cited alongside, same era.
Causal scrubbing: a method for rigorously testing interpretability hypotheses [redwood research]
Explaining grokking through circuit efficiency, 2023
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Later among the works it cites.
Neurons in large language models: Dead, n-gram, positional, 2023
Elena Voita, Javier Ferrando, and Christoforos Nalmpantis · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Qinan Yu, Jack Merullo, and Ellie Pavlick · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2023
Cited alongside, same era.
When is multilinguality a curse? language modeling for 250 high-and low-resource languages
Tyler A Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K Bergen · 2023
Cited alongside, same era.
How do languages influence each other? studying cross-lingual data sharing during llm fine-tuning
Rochelle Choenni, Dan Garrette, and Ekaterina Shutova · 2023
Cited alongside, same era.
Evaluating the ripple effects of knowledge editing in language models, 2023
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
Analyzing transformers in embedding space, 2023
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2023
Cited alongside, same era.
Multilingual jailbreak challenges in large language models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2023
Cited alongside, same era.
Causal abstraction for faithful model interpretation
Atticus Geiger, Christopher Potts, and Thomas Icard · 2023
Cited alongside, same era.
Later among the works it cites.
Aya 23: Open weight releases to further multilingual progress
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Kelly Marchisio, Sebastian Ruder, et al · 2024
Closest in time.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L Leavitt, and Naomi Saphra · 2024
Closest in time.
Examining modularity in multilingual lms via language-specialized subnetworks
Rochelle Choenni, Ekaterina Shutova, and Dan Garrette · 2024
Closest in time.
Information flow routes: Automatically interpreting language models at scale
Javier Ferrando and Elena Voita · 2024
Closest in time.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms, 2024
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov · 2024
Closest in time.
Preference tuning for toxicity mitigation generalizes across languages
Xiaochen Li, Zheng-Xin Yong, and Stephen H Bach · 2024
Closest in time.
Circuit component reuse across tasks in transformer language models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2024
Closest in time.
Kanishka Misra and Najoung Kim · 2024
Closest in time.
What needs to go right for an induction head? a mechanistic study of in-context learning circuits and their formation, 2024
Aaditya K. Singh, Ted Moskovitz, Felix Hill, Stephanie C. Y. Chan, and Andrew M. Saxe · 2024
Closest in time.
Language-specific neurons: The key to multilingual capabilities in large language models
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen · 2024
Closest in time.
Lm transparency tool: Interactive tool for analyzing transformer language models
Igor Tufanov, Karen Hambardzumyan, Javier Ferrando, and Elena Voita · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al · 2024
Closest in time.
Do llamas work in english? on the latent language of multilingual transformers
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West · 2024
Closest in time.
Reft: Representation finetuning for language models
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D Manning, and Christopher Potts · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.