Fetching the paper…
Reading the bibliography…
There is an ongoing debate in the NLP community whether modern language models contain linguistic knowledge, recovered through so-called probes.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
A synopsis of linguistic theory, 1930-1955
Firth, J. R. (1957) · 1955
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, M. P., Santorini, B., and Marcinkiewicz, M. A. (1993) · 1993
Earlier work this paper cites.
On structuring probabilistic dependences in stochastic language modelling
Ney, H., Essen, U., and Kneser, R. (1994) · 1994
Earlier work this paper cites.
Support-vector networks
Cortes, C., and Vapnik, V. (1995) · 1995
Earlier work this paper cites.
Placing search in context: the concept revisited
Finkelstein, L., Gabrilovich, E., Matias, Y., Rivlin, E., Solan, Z., Wolfman, G., and Ruppin, E. (2002) · 2002
Earlier work this paper cites.
A primer in bertology: What we know about how BERT works
Rogers, A., Kovaleva, O., and Rumshisky, A. (2020) · 2002
Earlier work this paper cites.
SRILM - an extensible language modeling toolkit
Stolcke, A. (2002) · 2002
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Bengio, Y., and Senecal, J. (2003) · 2003
Earlier work this paper cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Ravichander, A., Belinkov, Y., and Hovy, E. H. (2020) · 2005
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A. M. (2020) · 2005
Earlier work this paper cites.
Conditional entropy and mutual information
Press, W., Teukolsky, S., Vetterling, W., and Flannery, B. (2007) · 2007
Earlier work this paper cites.
Predicting what you already know helps: Provable self-supervised learning
Lee, J. D., Lei, Q., Saunshi, N., and Zhuo, J. (2020) · 2008
Earlier work this paper cites.
When do you need billions of words of pretraining data?
Zhang, Y., Warstadt, A., Li, H., and Bowman, S. R. (2020) · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013a) · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013b) · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Mikolov, T., Yih, W., and Zweig, G. (2013c) · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Mikolov, T., Yih, W., and Zweig, G. (2013d) · 2013
Earlier work this paper cites.
Ontonotes release 5.0
Weischedel, R., Pradhan, S., Ramshaw, L., Palmer, M., Xue, N., Marcus, M., Taylor, A., Greenberg, C., Hovy, E., Belvin, R., et al. (2013) · 2013
Earlier work this paper cites.
Neural word embedding as implicit matrix factorization
Levy, O., and Goldberg, Y. (2014) · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D. (2014) · 2014
Earlier work this paper cites.
A gold standard dependency corpus for english
Silveira, N., Dozat, T., de Marneffe, M., Bowman, S. R., Connor, M., Bauer, J., and Manning, C. D. (2014) · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Ganin, Y., and Lempitsky, V. S. (2015) · 2015
Earlier work this paper cites.
Visual information theory
Olah, C. (2015) · 2015
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. (2016) · 2016
Earlier work this paper cites.
The IWSLT 2016 evaluation campaign
Cettolo, M., Jan, N., Sebastian, S., Bentivogli, L., Cattoni, R., and Federico, M. (2016) · 2016
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Ettinger, A., Elgohary, A., and Resnik, P. (2016) · 2016
Earlier work this paper cites.
Word embeddings as metric recovery in semantic spaces
Hashimoto, T. B., Alvarez-Melis, D., and Jaakkola, T. S. (2016) · 2016
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Linzen, T., Dupoux, E., and Goldberg, Y. (2016) · 2016
Cited alongside, same era.
Does string-based neural MT learn source syntax?
Shi, X., Padhi, I., and Knight, K. (2016) · 2016
Cited alongside, same era.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Adi, Y., Kermany, E., Belinkov, Y., Lavi, O., and Goldberg, Y. (2017) · 2017
Cited alongside, same era.
Skip-gram - zipf + uniform = vector additivity
Gittens, A., Achlioptas, D., and Mahoney, M. W. (2017) · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, H., Khosravi, K., and Socher, R. (2017) · 2017
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Klein, G., Kim, Y., Deng, Y., Senellart, J., and Rush, A. M. (2017) · 2017
What do you learn from context? probing for sentence structure in contextualized word representations
Tenney, I., Xia, P., Chen, B., Wang, A., Poliak, A., McCoy, R. T., Kim, N., Durme, B. V., Bowman, S. R., Das, D., and Pavlick, E. (2019b) · 2019
Later among the works it cites.
Near-lossless binarization of word embeddings
Tissier, J., Gravier, C., and Habrard, A. (2019) · 2019
Later among the works it cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Bender, E. M., and Koller, A. (2020) · 2020
Later among the works it cites.
The lottery ticket hypothesis for pre-trained BERT networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M. (2020) · 2020
Later among the works it cites.
How to probe sentence embeddings in low-resource languages: On structural design choices for probing task evaluation
Eger, S., Daxenberger, J., and Gurevych, I. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learned in translation: Contextualized word vectors
McCann, B., Bradbury, J., Xiong, C., and Socher, R. (2017) · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2017) · 2017
Cited alongside, same era.
The mechanism of additive composition
Tian, R., Okazaki, N., and Inui, K. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Evaluating the stability of embedding-based word similarities
Antoniak, M., and Mimno, D. (2018) · 2018
Cited alongside, same era.
What you can cram into a single \$&!#* vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M. (2018) · 2018
Cited alongside, same era.
Compressing BERT: studying the effects of weight pruning on transfer learning
Gordon, M. A., Duh, K., and Andrews, N. (2020) · 2020
Later among the works it cites.
A mutual information maximization perspective of language representation learning
Kong, L., de Masson d’Autume, C., Yu, L., Ling, W., Dai, Z., and Yogatama, D. (2020) · 2020
Later among the works it cites.
Classifier probes may just learn from linear context features
Kunz, J., and Kuhlmann, M. (2020) · 2020
Later among the works it cites.
A tale of a probe and a parser
Maudslay, R. H., Valvoda, J., Pimentel, T., Williams, A., and Cotterell, R. (2020) · 2020
Later among the works it cites.
Asking without telling: Exploring latent ontologies in contextual representations
Michael, J., Botha, J. A., and Tenney, I. (2020) · 2020
Later among the works it cites.
Pareto probing: Trading off accuracy for complexity
Pimentel, T., Saphra, N., Williams, A., and Cotterell, R. (2020a) · 2020
Later among the works it cites.
Information-theoretic probing for linguistic structure
Pimentel, T., Valvoda, J., Maudslay, R. H., Zmigrod, R., Williams, A., and Cotterell, R. (2020b) · 2020
Later among the works it cites.
When BERT plays the lottery, all tickets are winning
Prasanna, S., Rogers, A., and Rumshisky, A. (2020) · 2020
Later among the works it cites.
jiant: A software toolkit for research on general-purpose text understanding models
Pruksachatkun, Y., Yeres, P., Liu, H., Phang, J., Htut, P. M., Wang, A., Tenney, I., and Bowman, S. R. (2020) · 2020
Later among the works it cites.
Stanza: A python natural language processing toolkit for many human languages
Qi, P., Zhang, Y., Zhang, Y., Bolton, J., and Manning, C. D. (2020) · 2020
Later among the works it cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y. (2020) · 2020
Later among the works it cites.
Investigating transferability in pretrained language models
Tamkin, A., Singh, T., Giovanardi, D., and Goodman, N. D. (2020) · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Voita, E., and Titov, I. (2020) · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. (2020) · 2020
Later among the works it cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Wu, Z., Chen, Y., Kao, B., and Liu, Q. (2020) · 2020
Later among the works it cites.
Masking as an efficient alternative to finetuning for pretrained language models
Zhao, M., Lin, T., Mi, F., Jaggi, M., and Schütze, H. (2020) · 2020
Later among the works it cites.
An information theoretic view on selecting linguistic probes
Zhu, Z., and Rudzicz, F. (2020) · 2020
Later among the works it cites.
Sgns implementation in pytorch
Assylbekov, Z. (2020) · 2021
Closest in time.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y. (2021) · 2021
Closest in time.
structural-probes
Hewitt, J. (2019) · 2021
Closest in time.
About the test data
Mahoney, M. (2011) · 2021
Closest in time.
Pretraining RoBERTa using your own data
Ott, M., Baevski, A., Martin, L., and Morton, J. (2019a) · 2021
Closest in time.
The singleton fallacy: Why current critiques of language models miss the point
Sahlgren, M., and Carlsson, F. (2021) · 2021
Closest in time.
A mathematical exploration of why language models help solve downstream tasks
Saunshi, N., Malladi, S., and Arora, S. (2021) · 2021
Closest in time.