Fetching the paper…
Reading the bibliography…
How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023).
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013c · 2013
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018 · 2018
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Earlier work this paper cites.
Residual connections encourage iterative inference
Stanisław Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio. 2018 · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2019 · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. 2020 · 2020
Earlier work this paper cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020 · 2020
Earlier work this paper cites.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas. 2020 · 2020
Cited alongside, same era.
interpreting gpt: the logit lens
nostalgebraist. 2020 · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020 · 2020
Cited alongside, same era.
Pareto probing: Trading off accuracy for complexity
Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020a · 2020
Cited alongside, same era.
On the pitfalls of analyzing individual neurons in language models
Omer Antverg and Yonatan Belinkov. 2021 · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Evaluation beyond task performance: Analyzing concepts in alphazero in hex
Charles Lovering, Jessica Forde, George Konidaris, Ellie Pavlick, and Michael Littman. 2022 · 2022
Later among the works it cites.
Acquisition of chess knowledge in alphazero
Thomas McGrath, Andrei Kapishnikov, Nenad Tomašev, Adam Pearce, Martin Wattenberg, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kramnik. 2022 · 2022
Later among the works it cites.
Linearly mapping from image to text space
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick. 2022 · 2022
Later among the works it cites.
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick. 2022 · 2022
Later among the works it cites.
Polysemanticity and capacity in neural networks
Adam Scherlis, Kshitij Sachan, Adam S Jermyn, Joe Benton, and Buck Shlegeris. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021 · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Belinda Z. Li, Maxwell Nye, and Jacob Andreas. 2021 · 2021
Cited alongside, same era.
What if this modified that? syntactic interventions with counterfactual embeddings
Mycal Tucker, Peng Qian, and Roger Levy. 2021 · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah. 2022a · 2022
Cited alongside, same era.
Extracting latent steering vectors from pretrained language models
Nishant Subramani, Nivedita Suresh, and Matthew Peters. 2022 · 2022
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023 · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023 · 2023
Closest in time.
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023 · 2023
Closest in time.
Measuring and manipulating knowledge representations in language models
Evan Hernandez, Belinda Z Li, and Jacob Andreas. 2023 · 2023
Closest in time.
The hydra effect: Emergent self-repair in language model computations
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. 2023 · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt. 2023 · 2023
Closest in time.
Steering gpt-2-xl by adding an activation vector - ai alignment forum
Alex Turner, Monte MacDiarmid, David Udell, lisathiergart, and Ulisse Mini. 2023 · 2023
Closest in time.