Fetching the paper…
Reading the bibliography…
There has been a lot of interest in understanding what information is captured by hidden representations of language models (LMs).
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams. 1992 · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Direct and Indirect Effects
Judea Pearl. 2001 · 2001
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe. 2004 · 2004
Earlier work this paper cites.
Inductive biases for deep learning of higher-level cognition
Anirudh Goyal and Yoshua Bengio. 2020 · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. 2014 · 2014
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai. 2016 · 2016
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017 · 2017
Earlier work this paper cites.
Investigating ‘Aspect’in NMT and SMT: Translating the English simple past and present perfect
Eva Vanmassenhove, Jinhua Du, and Andy Way. 2017 · 2017
Cited alongside, same era.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Learning sparse neural networks through l_0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Cited alongside, same era.
Gender Bias in Neural Natural Language Processing
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2020 · 2020
Later among the works it cites.
Visualizing the Impact of Feature Attribution Baselines
Pascal Sturmfels, Scott Lundberg, and Su-In Lee. 2020 · 2020
Later among the works it cites.
Investigating transferability in pretrained language models
Alex Tamkin, Trisha Singh, Davide Giovanardi, and Noah Goodman. 2020 · 2020
Later among the works it cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020 · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Later among the works it cites.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What do RNN language models learn about filler–gap dependencies?
Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell. 2018 · 2018
Cited alongside, same era.
Adversarial removal of demographic attributes revisited
Maria Barrett, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, and Anders Søgaard. 2019 · 2019
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Cited alongside, same era.
Analysing neural language models: Contextual decomposition reveals default reasoning in number and gender assignment
Jaap Jumelet, Willem Zuidema, and Dieuwke Hupkes. 2019 · 2019
Cited alongside, same era.
The emergence of number and syntax units in LSTM language models
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber. 2021 · 2021
Closest in time.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2021 · 2021
Closest in time.
On the transfer of disentangled representations in realistic settings
Andrea Dittadi, Frederik Träuble, Francesco Locatello, Manuel Wuthrich, Vaibhav Agrawal, Ole Winther, Stefan Bauer, and Bernhard Schölkopf. 2021 · 2021
Closest in time.
Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Closest in time.
CausaLM: Causal Model Explanation Through Counterfactual Language Models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2021 · 2021
Closest in time.
Language models use monotonicity to assess NPI licensing
Jaap Jumelet, Milica Denic, Jakub Szymanik, Dieuwke Hupkes, and Shane Steinert-Threlkeld. 2021 · 2021
Closest in time.
Mechanisms for handling nested dependencies in neural-network language models and humans
Yair Lakretz, Dieuwke Hupkes, Alessandra Vergallito, Marco Marelli, Marco Baroni, and Stanislas Dehaene. 2021 · 2021
Closest in time.
Interpreting graph neural networks for NLP with differentiable edge masking
Michael Sejr Schlichtkrull, Nicola De Cao, and Ivan Titov. 2021 · 2021
Closest in time.