Fetching the paper…
Reading the bibliography…
Causal probing aims to analyze foundation models by examining how intervening on their representation of various latent properties impacts their outputs.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and 1 others. 2019 · 1910
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A Paszke. 2019 · 1912
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber. 2020 · 2004
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, and 1 others. 2019 · 2019
Earlier work this paper cites.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Francesco Croce and Matthias Hein. 2020 · 2020
Earlier work this paper cites.
Neural natural language inference models partially embed theories of lexical entailment and negation
Atticus Geiger, Kyle Richardson, and Christopher Potts. 2020 · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Exploring the linear subspace hypothesis in gender bias mitigation
Francisco Vargas and Ryan Cotterell. 2020 · 2020
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction
Shauli Ravfogel, Grusha Prasad, Tal Linzen, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Competence-based analysis of language models
Adam Davies, Jize Jiang, and ChengXiang Zhai. 2023 · 2023
Later among the works it cites.
Log-linear guardedness and its implications
Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023 · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2021 · 2021
Cited alongside, same era.
What if this modified that? syntactic interventions with counterfactual embeddings
Mycal Tucker, Peng Qian, and Roger Levy. 2021 · 2021
Cited alongside, same era.
A non-linear structural probe
Jennifer C. White, Tiago Pimentel, Naomi Saphra, and Ryan Cotterell. 2021 · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Probing classifiers are unreliable for concept removal and detection
Abhinav Kumar, Chenhao Tan, and Amit Sharma. 2022 · 2022
Cited alongside, same era.
Probing for the usage of grammatical number
Karim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau, and Ryan Cotterell. 2022 · 2022
Cited alongside, same era.
Adversarial concept erasure in kernel space
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg, and Ryan Cotterell. 2022b · 2022
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, and 1 others. 2023 · 2023
Later among the works it cites.
Causalgym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman Arora, Dan Jurafsky, and Christopher Potts. 2024 · 2024
Closest in time.
Adam Davies and Ashkan Khakzar. 2024 · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah Goodman. 2024 · 2024
Closest in time.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024 · 2024
Closest in time.
RAVEL: Evaluating interpretability methods on disentangling language model representations
Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, and Atticus Geiger. 2024 · 2024
Closest in time.
On the origins of linear representations in large language models
Yibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam, and Victor Veitch. 2024 · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda. 2024 · 2024
Closest in time.