Fetching the paper…
Reading the bibliography…
Concept erasure aims to remove specified features from an embedding.
Mathematical Statistics
Thomas S. Ferguson · 1967
Earlier work this paper cites.
Making Things Happen: A Theory of Causal Explanation explanation
James Francis Woodward · 2005
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Universal dependency annotation for multilingual parsing
Ryan McDonald, Joakim Nivre, Yvonne Quirmbach-Brundage, Yoav Goldberg, Dipanjan Das, Kuzman Ganchev, Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, et al · 2013
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam T. Kalai · 2016
Earlier work this paper cites.
Censoring representations with an adversary
Harrison Edwards and Amos Storkey · 2016
Earlier work this paper cites.
Counterfactual fairness
Matt J. Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva · 2017
Earlier work this paper cites.
Controllable invariance through adversarial feature learning
Qizhe Xie, Zihang Dai, Yulun Du, Eduard Hovy, and Graham Neubig · 2017
Earlier work this paper cites.
The Hilbert space of random variables
UC Berkeley · 2018
Earlier work this paper cites.
Adversarial deep averaging networks for cross-lingual sentiment classification
Xilun Chen, Yu Sun, Ben Athiwaratkun, Claire Cardie, and Kilian Weinberger · 2018
Earlier work this paper cites.
Adversarial removal of demographic attributes from text data
Yanai Elazar and Yoav Goldberg · 2018
Earlier work this paper cites.
Mitigating unwanted biases with adversarial learning
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
On the global optima of kernelized adversarial representation learning
Bashir Sadeghi, Runyi Yu, and Vishnu Boddeti · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Why attention is not explanation: Surgical intervention and causal reasoning about neural models
Christopher Grimsley, Elijah Mayfield, and Julia R.S. Bursten · 2020
Cited alongside, same era.
spaCy: Industrial-strength Natural Language Processing in Python, 2020
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd · 2020
Cited alongside, same era.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Can transformer be too compositional? Analysing idiom processing in neural machine translation
Verna Dankers, Christopher Lucas, and Ivan Titov · 2022
Later among the works it cites.
Inducing causal structure for interpretable neural networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts · 2022
Later among the works it cites.
Better hit the nail on the head than beat around the bush: Removing protected attributes with a single projection
Pantea Haghighatkhah, Antske Fokkens, Pia Sommerauer, Bettina Speckmann, and Kevin Verbeek · 2022
Later among the works it cites.
Probing classifiers are unreliable for concept removal and detection
Abhinav Kumar, Chenhao Tan, and Amit Sharma · 2022
Later among the works it cites.
Causal conceptions of fairness and their consequences
Hamed Nilforoshan, Johann D. Gaebler, Ravi Shroff, and Sharad Goel · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon · 2020
Cited alongside, same era.
OSCaR: Orthogonal subspace correction and rectification of biases in word embeddings
Sunipa Dev, Tao Li, Jeff M. Phillips, and Vivek Srikumar · 2021
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Hatfield Zac Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Obstructing classification via projection
Pantea Haghighatkhah, Wouter Meulemans, Bettina Speckmann, Jérôme Urhausen, and Kevin Verbeek · 2021
Cited alongside, same era.
The low-dimensional linear geometry of contextualized word representations
Evan Hernandez and Jacob Andreas · 2021
Cited alongside, same era.
The rediscovery hypothesis: Language models need to meet linguistics
Vassilina Nikoulina, Maxat Tezekbayev, Nuradil Kozhakhmet, Madina Babazhanova, Matthias Gallé, and Zhenisbek Assylbekov · 2021
Cited alongside, same era.
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg, and Ryan Cotterell · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
Efficient fair PCA for fair representation learning
Matthäus Kleindessner, Michele Donini, Chris Russell, and Muhammad Bilal Zafar · 2023
Closest in time.
Actually, Othello-GPT has a linear emergent world model, Mar 2023
Neel Nanda · 2023
Closest in time.
Log-linear guardedness and its implications
Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell · 2023
Closest in time.
Gold doesn’t always glitter: Spectral removal of linear and nonlinear guarded attribute information
Shun Shao, Yftah Ziser, and Shay B. Cohen · 2023
Closest in time.
Erasure of unaligned attributes from neural representations
Shun Shao, Yftah Ziser, and Shay B. Cohen · 2023
Closest in time.
LLaMA: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Interpretability at scale: Identifying causal mechanisms in Alpaca
Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah Goodman · 2023
Closest in time.