Fetching the paper…
Reading the bibliography…
The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information.
Assessing BERT’s Syntactic Abilities
Goldberg, Y. 2019 · 1901
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; and Brew, J. 2020 · 1910
Earlier work this paper cites.
Asymptotic Theory of Rejective Sampling with Varying Probabilities from a Finite Population
Hájek, J. 1964 · 1964
Earlier work this paper cites.
Probing Linguistic Systematicity
Goodwin, E.; Sinha, K.; and O’Donnell, T. J. 2020 · 1969
Earlier work this paper cites.
A Simple Sequentially Rejective Multiple Test Procedure
Holm, S. 1979 · 1979
Earlier work this paper cites.
An Estimate of an Upper Bound for the Entropy of English
Brown, P. F.; Della Pietra, S. A.; Della Pietra, V. J.; Lai, J. C.; and Mercer, R. L. 1992 · 1992
Earlier work this paper cites.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
Williams, R. J. 1992 · 1992
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Algorithms to find exact inclusion probabilities for conditional Poisson sampling and Pareto π \pi s sampling designs
Aires, N. 1999 · 1999
Earlier work this paper cites.
Parameter Estimation for Probabilistic Finite-State Transducers
Eisner, J. 2002 · 2002
Earlier work this paper cites.
Regularization and variable selection via the Elastic Net
Zou, H.; and Hastie, T. 2005 · 2005
Earlier work this paper cites.
First- and Second-Order Expectation Semirings with Applications to Minimum-Risk Training on Translation Forests
Li, Z.; and Eisner, J. 2009 · 2009
Earlier work this paper cites.
Rectified Linear Units Improve Restricted Boltzmann Machines
Nair, V.; and Hinton, G. E. 2010 · 2010
Earlier work this paper cites.
Determinantal Point Processes for Machine Learning
Kulesza, A. 2012 · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Murphy, K. P. 2012 · 2012
Earlier work this paper cites.
Stochastic variational inference
Hoffman, M. D.; Blei, D. M.; Wang, C.; and Paisley, J. 2013 · 2013
Earlier work this paper cites.
Visualizing and Understanding Recurrent Networks
Karpathy, A.; Johnson, J.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Visualizing and Understanding Neural Models in NLP
Li, J.; Chen, X.; Hovy, E.; and Jurafsky, D. 2016 · 2016
Earlier work this paper cites.
Understanding Neural Networks through Representation Erasure
Li, J.; Monroe, W.; and Jurafsky, D. 2016 · 2016
Earlier work this paper cites.
Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
Linzen, T.; Dupoux, E.; and Goldberg, Y. 2016 · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Cited alongside, same era.
Representation of Linguistic Form and Function in Recurrent Neural Networks
Kádár, Á.; Chrupała, G.; and Alishahi, A. 2017 · 2017
Cited alongside, same era.
Universal Dependencies 2.1
Nivre, J.; Agić, Ž.; Ahrenberg, L.; Antonsen, L.; Aranzabe, M. J.; Asahara, M.; Ateyah, L.; Attia, M.; Atutxa, A.; Augustinus, L.; Badmaeva, E.; Ballesteros, M.; Banerjee, E.; Bank, S.; Barbu Mititelu, V.; Bauer, J.; Bengoetxea, K.; Bhat, R. A.; Bick, E.; Bobicev, V.; Börstell, C.; Bosco, C.; Bouma, G.; Bowman, S.; Burchardt, A.; Candito, M.; Caron, G.; Cebiroğlu Eryiğit, G.; Celano, G. G. A.; Cetin, S.; Chalub, F.; Choi, J.; Cinková, S.; Çöltekin, Ç.; Connor, M.; Davidson, E.; de Marneffe, M.-C.; de Paiva, V.; Diaz de Ilarraza, A.; Dirix, P.; Dobrovoljc, K.; Dozat, T.; Droganova, K.; Dwivedi, P.; Eli, M.; Elkahky, A.; Erjavec, T.; Farkas, R.; Fernandez Alcalde, H.; Foster, J.; Freitas, C.; Gajdošová, K.; Galbraith, D.; Garcia, M.; Gärdenfors, M.; Gerdes, K.; Ginter, F.; Goenaga, I.; Gojenola, K.; Gökırmak, M.; Goldberg, Y.; Gómez Guinovart, X.; Gonzáles Saavedra, B.; Grioni, M.; Grūzītis, N.; Guillaume, B.; Habash, N.; Hajič, J.; Hajič jr., J.; Hà Mỹ, L.; Harris, K.; Haug, D.; Hladká, B.; Hlaváčová, J.; Hociung, F.; Hohle, P.; Ion, R.; Irimia, E.; Jelínek, T.; Johannsen, A.; Jørgensen, F.; Kaşıkara, H.; Kanayama, H.; Kanerva, J.; Kayadelen, T.; Kettnerová, V.; Kirchner, J.; Kotsyba, N.; Krek, S.; Laippala, V.; Lambertino, L.; Lando, T.; Lee, J.; Lê Hồng, P.; Lenci, A.; Lertpradit, S.; Leung, H.; Li, C. Y.; Li, J.; Li, K.; Ljubešić, N.; Loginova, O.; Lyashevskaya, O.; Lynn, T.; Macketanz, V.; Makazhanov, A.; Mandl, M.; Manning, C.; Mărănduc, C.; Mareček, D.; Marheinecke, K.; Martínez Alonso, H.; Martins, A.; Mašek, J.; Matsumoto, Y.; McDonald, R.; Mendonça, G.; Miekka, N.; Missilä, A.; Mititelu, C.; Miyao, Y.; Montemagni, S.; More, A.; Moreno Romero, L.; Mori, S.; Moskalevskyi, B.; Muischnek, K.; Müürisep, K.; Nainwani, P.; Nedoluzhko, A.; Nešpore-Bērzkalne, G.; Nguyễn Thị, L.; Nguyễn Thị Minh, H.; Nikolaev, V.; Nurmi, H.; Ojala, S.; Osenova, P.; Östling, R.; Øvrelid, L.; Pascual, E.; Passarotti, M.; Perez, C.-A.; and Perrier, G. e. a. 2017 · 2017
Understanding the Role of Individual Units in a Deep Neural Network
Bau, D.; Zhu, J.-Y.; Strobelt, H.; Lapedriza, A.; Zhou, B.; and Torralba, A. 2020 · 2020
Later among the works it cites.
Analyzing Individual Neurons in Pre-trained Language Models
Durrani, N.; Sajjad, H.; Dalvi, F.; and Belinkov, Y. 2020 · 2020
Later among the works it cites.
A Tale of a Probe and a Parser
Hall Maudslay, R.; Valvoda, J.; Pimentel, T.; Williams, A.; and Cotterell, R. 2020 · 2020
Later among the works it cites.
Compositional Explanations of Neurons
Mu, J.; and Andreas, J. 2020 · 2020
Later among the works it cites.
Information-Theoretic Probing for Linguistic Structure
Pimentel, T.; Valvoda, J.; Hall Maudslay, R.; Zmigrod, R.; Williams, A.; and Cotterell, R. 2020 · 2020
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
What You Can Cram into a Single $&!#* Vector: Probing Sentence Embeddings for Linguistic Properties
Conneau, A.; Kruszewski, G.; Lample, G.; Barrault, L.; and Baroni, M. 2018 · 2018
Cited alongside, same era.
Under the Hood: Using Diagnostic Classifiers to Investigate and Improve How Language Models Track Agreement Information
Giulianelli, M.; Harding, J.; Mohnert, F.; Hupkes, D.; and Zuidema, W. 2018 · 2018
Cited alongside, same era.
Colorless Green Recurrent Networks Dream Hierarchically
Gulordava, K.; Bojanowski, P.; Grave, E.; Linzen, T.; and Baroni, M. 2018 · 2018
Cited alongside, same era.
Marrying Universal Dependencies and Universal Morphology
McCarthy, A. D.; Silfverberg, M.; Cotterell, R.; Hulden, M.; and Yarowsky, D. 2018 · 2018
Cited alongside, same era.
Deep Contextualized Word Representations
Peters, M.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Cited alongside, same era.
Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation
Poliak, A.; Haldar, A.; Rudinger, R.; Hu, J. E.; Pavlick, E.; White, A. S.; and Van Durme, B. 2018 · 2018
Cited alongside, same era.
What Do You Learn from Context? Probing for Sentence Structure in Contextualized Word Representations
Tenney, I.; Xia, P.; Chen, B.; Wang, A.; Poliak, A.; McCoy, R. T.; Kim, N.; Durme, B. V.; Bowman, S. R.; Das, D.; and Pavlick, E. 2018 · 2018
Cited alongside, same era.
Language Modeling Teaches You More than Translation Does: Lessons Learned Through Auxiliary Syntactic Task Analysis
Zhang, K.; and Bowman, S. 2018 · 2018
Cited alongside, same era.
Identifying and Controlling Important Neurons in Neural Machine Translation
Bau, A.; Belinkov, Y.; Sajjad, H.; Durrani, N.; Dalvi, F.; and Glass, J. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
TX-Ray: Quantifying and explaining model-knowledge transfer in (un-)supervised NLP
Rethmeier, N.; Saxena, V. K.; and Augenstein, I. 2020 · 2020
Later among the works it cites.
A Primer in BERTology: What We Know About How BERT Works
Rogers, A.; Kovaleva, O.; and Rumshisky, A. 2020 · 2020
Later among the works it cites.
Torch-Struct: Deep Structured Prediction Library
Rush, A. 2020 · 2020
Later among the works it cites.
LINSPECTOR: Multilingual Probing Tasks for Word Representations
Şahin, G. G.; Vania, C.; Kuznetsov, I.; and Gurevych, I. 2020 · 2020
Later among the works it cites.
Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English
Tang, G.; Sennrich, R.; and Nivre, J. 2020 · 2020
Later among the works it cites.
Intrinsic Probing through Dimension Selection
Torroba Hennigen, L.; Williams, A.; and Cotterell, R. 2020 · 2020
Later among the works it cites.
Investigating Gender Bias in Language Models Using Causal Mediation Analysis
Vig, J.; Gehrmann, S.; Belinkov, Y.; Qian, S.; Nevo, D.; Singer, Y.; and Shieber, S. 2020 · 2020
Later among the works it cites.
Information-Theoretic Probing with Minimum Description Length
Voita, E.; and Titov, I. 2020 · 2020
Later among the works it cites.
Probing Pretrained Language Models for Lexical Semantics
Vulić, I.; Ponti, E. M.; Litschko, R.; Glavaš, G.; and Korhonen, A. 2020 · 2020
Later among the works it cites.
Subword Pooling Makes a Difference
Ács, J.; Kádár, Á.; and Kornai, A. 2021 · 2021
Later among the works it cites.
On the Pitfalls of Analyzing Individual Neurons in Language Models
Antverg, O.; and Belinkov, Y. 2021 · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and alternatives
Belinkov, Y. 2021 · 2021
Later among the works it cites.
Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?
Ravichander, A.; Belinkov, Y.; and Hovy, E. 2021 · 2021
Later among the works it cites.
Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Geva, M.; Caciularu, A.; Wang, K. R.; and Goldberg, Y. 2022 · 2022
Closest in time.
Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models
Stańczak, K.; Ponti, E.; Torroba Hennigen, L.; Cotterell, R.; and Augenstein, I. 2022 · 2022
Closest in time.