Fetching the paper…
Reading the bibliography…
When language models process syntactically complex sentences, do they use their representations of syntax in a manner that is consistent with the grammar of the language? We propose AlterRep, an intervention-based method to address this question.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Causal mediation analysis for interpreting neural NLP: the case of gender bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020 · 2004
Earlier work this paper cites.
CausaLM: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2020 · 2005
Earlier work this paper cites.
Making things happen: A theory of causal explanation
James Woodward. 2005 · 2005
Earlier work this paper cites.
Counterfactual explanations for machine learning: A review
Sahil Verma, John P. Dickerson, and Keegan Hines. 2020 · 2010
Earlier work this paper cites.
Analyzing the source and target contributions to predictions in neural machine translation
Elena Voita, Rico Sennrich, and Ivan Titov. 2020 · 2010
Earlier work this paper cites.
Explaining NLP models via minimal contrastive editing (mice)
Alexis Ross, Ana Marasovic, and Matthew E. Peters. 2020 · 2012
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, Germán Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem H. Zuidema. 2018 · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Contrastive explanation: A structural-model approach
Tim Miller. 2018 · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Cited alongside, same era.
The emergence of number and syntax units in LSTM language models
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni. 2019 · 2019
Cited alongside, same era.
It’s all in the name: Mitigating gender bias with name-based counterfactual data substitution
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Neutralizing gender bias in word embedding with latent disentanglement and counterfactual generation
Seungjae Shin, Kyungwoo Song, JoonHo Jang, Hyemi Kim, Weonyoung Joo, and Il-Chul Moon. 2020 · 2020
Later among the works it cites.
Investigating transferability in pretrained language models
Alex Tamkin, Trisha Singh, Davide Giovanardi, and Noah D. Goodman. 2020 · 2020
Later among the works it cites.
Blimp: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2020 · 2020
Later among the works it cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rowan Hall Maudslay, Hila Gonen, Ryan Cotterell, and Simone Teufel. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2019 · 2019
Cited alongside, same era.
Using priming to uncover the organization of syntactic representations in neural language models
Grusha Prasad, Marten van Schijndel, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Ran Zmigrod, S. J. Mielke, Hanna M. Wallach, and Ryan Cotterell. 2019 · 2019
Cited alongside, same era.
Neural language models capture some, but not all, agreement attraction effects
Suhas Arehalli and Tal Linzen. 2020 · 2020
Cited alongside, same era.
Syntaxgym: An online platform for targeted evaluation of language models
Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy. 2020 · 2020
Cited alongside, same era.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov. 2021 · 2021
Closest in time.
ECINN: efficient counterfactuals from invertible neural networks
Frederik Hvilshøj, Alexandros Iosifidis, and Ira Assent. 2021 · 2021
Closest in time.
Contrastive explanations for model interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Closest in time.
Mechanisms for handling nested dependencies in neural-network language models and humans
Yair Lakretz, Dieuwke Hupkes, Alessandra Vergallito, Marco Marelli, Marco Baroni, and Stanislas Dehaene. 2021 · 2021
Closest in time.
Causal effects of linguistic properties
Reid Pryzant, Dallas Card, Dan Jurafsky, Victor Veitch, and Dhanya Sridhar. 2021 · 2021
Closest in time.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard H. Hovy. 2021 · 2021
Closest in time.
Mediators in determining what processing BERT performs first
Aviv Slobodkin, Leshem Choshen, and Omri Abend. 2021 · 2021
Closest in time.
What if this modified that? syntactic interventions via counterfactual embeddings
Mycal Tucker, Peng Qian, and Roger Levy. 2021 · 2021
Closest in time.