Fetching the paper…
Reading the bibliography…
The linear subspace hypothesis (Bolukbasi et al., 2016) states that, in a language model's representation space, all information about a concept such as verbal number is encoded in a linear subspace.
A mathematical theory of communication
Claude E. Shannon. 1948 · 1948
Earlier work this paper cites.
Elements of Information Theory , second edition
Thomas M. Cover and Joy M. Thomas. 2006 · 2006
Earlier work this paper cites.
Syncretism
Matthew Baerman. 2007 · 2007
Earlier work this paper cites.
Causal inference in statistics: An overview
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Towards a Universal Stanford Dependencies parallel treebank
Cristina Bosco and Manuela Sanguinetti. 2014 · 2014
Earlier work this paper cites.
Rhapsodie: a prosodic-syntactic treebank for spoken French
Anne Lacheret, Sylvain Kahane, Julie Beliao, Anne Dister, Kim Gerdes, Jean-Philippe Goldman, Nicolas Obin, Paola Pietrandrea, and Atanas Tchobanov. 2014 · 2014
Earlier work this paper cites.
Converting the parallel treebank ParTUT in Universal Stanford Dependencies
Manuela Sanguinetti and Cristina Bosco. 2014 · 2014
Earlier work this paper cites.
SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation
Felix Hill, Roi Reichart, and Anna Korhonen. 2015 · 2015
Earlier work this paper cites.
PartTUT: The Turin University Parallel Treebank , chapter 1. Springer International Publishing, Cham
Manuela Sanguinetti and Cristina Bosco. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam T. Kalai. 2016 · 2016
Earlier work this paper cites.
Generating sentences from a continuous space
Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew Dai, Rafal Jozefowicz, and Samy Bengio. 2016 · 2016
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Earlier work this paper cites.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 2019
Earlier work this paper cites.
Conversion et améliorations de corpus du français annotés en Universal Dependencies
Bruno Guillaume, Marie-Catherine de Marneffe, and Guy Perrier. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. 2020 · 2020
spaCy: Industrial-strength natural language processing in Python
Ines Montani, Matthew Honnibal, Matthew Honnibal, Sofie Van Landeghem, Adriane Boyd, Henning Peters, Paul O’Leary McCann, Maxim Samsonov, Jim Geovedi, Jim O’Regan, Duygu Altinok, György Orosz, Søren Lind Kristiansen, , Roman, Explosion Bot, Lj Miranda, Leander Fiedler, Daniël De Kok, Grégory Howard, , Edward, Wannaphong Phatthiyaphaibun, Yohei Tamura, Sam Bozek, , Murat, Mark Amery, Ryn Daniels, Björn Böing, Pradeep Kumar Tippa, and Peter Baumgartner. 2022 · 2022
Later among the works it cites.
Introducing ChatGPT
OpenAI. 2022 · 2022
Later among the works it cites.
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg, and Ryan Cotterell. 2022b · 2022
Later among the works it cites.
Naturalistic causal probing for morpho-syntax
Afra Amini, Tiago Pimentel, Clara Meister, and Ryan Cotterell. 2023 · 2023
Closest in time.
LEACE: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
CausaLM: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2021 · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021 · 2021
Cited alongside, same era.
Contrastive explanations for model interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction
Shauli Ravfogel, Grusha Prasad, Tal Linzen, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein. 2021 · 2021
Cited alongside, same era.
CEBaB: Estimating the causal effects of real-world concepts on NLP model behavior
Eldar David Abraham, Karel D’Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu. 2022 · 2022
Cited alongside, same era.
Closest in time.
A measure-theoretic characterization of tight language models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell. 2023 · 2023
Closest in time.
Wikimedia downloads
Wikimedia Foundation. 2023 · 2023
Closest in time.
Causal abstraction for faithful model interpretation
Atticus Geiger, Chris Potts, and Thomas Icard. 2023 · 2023
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch. 2023 · 2023
Closest in time.
Log-linear guardedness and its implications
Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Closest in time.
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman Arora, Dan Jurafsky, and Christopher Potts. 2024 · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah Goodman. 2024 · 2024
Closest in time.
ReFT: Representation finetuning for language models
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D. Manning, and Christopher Potts. 2024 · 2024
Closest in time.