Fetching the paper…
Reading the bibliography…
We study how masking and predicting tokens in an unsupervised fashion can give rise to linguistic structures and downstream performance gains.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
A. Wang and K. Cho. 2019 · 1902
Earlier work this paper cites.
Fine-tune BERT for extractive summarization
Y. Liu. 2019 · 1903
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. 2019b · 1907
Earlier work this paper cites.
Matrix Perturbation Theory
G. W. Stewart and J. Sun. 1990 · 1990
Earlier work this paper cites.
Two experiments on learning probabilistic dependency grammars from corpora
G. Carroll and E. Charniak. 1992 · 1992
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani. 1996 · 1996
Earlier work this paper cites.
Thumbs up? sentiment classification using machine learning techniques
B. Pang, L. Lee, and S. Vaithyanathan. 2002 · 2002
Earlier work this paper cites.
Grammatical bigrams
M. A. Paskin. 2002 · 2002
Earlier work this paper cites.
Latent Dirichlet allocation
D. Blei, A. Ng, and M. I. Jordan. 2003 · 2003
Earlier work this paper cites.
Mining and summarizing customer reviews
M. Hu and B. Liu. 2004 · 2004
Earlier work this paper cites.
Corpus-based induction of syntactic structure: Models of dependency and constituency
D. Klein and C.D. Manning. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. 2020 · 2005
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
M. de Marneffe, B. MacCartney, and C. D. Manning. 2006 · 2006
Earlier work this paper cites.
High-dimensional graphs and variable selection with the lasso
N. Meinshausen and P. Bühlmann. 2006 · 2006
Earlier work this paper cites.
Revisiting few-sample bert fine-tuning
T. Zhang, F. Wu, A. Katiyar, K. Q. Weinberger, and Y. Artzi. 2020 · 2006
Earlier work this paper cites.
Predicting what you already know helps: Provable self-supervised learning
J. D. Lee, Q. Lei, N. Saunshi, and J. Zhuo. 2020 · 2008
Earlier work this paper cites.
A mathematical exploration of why language models help solve downstream tasks
N. Saunshi, S. Malladi, and S. Arora. 2020 · 2010
Cited alongside, same era.
Second-order unsupervised neural dependency parsing
S. Yang, Y. Jiang, W. Han, and K. Tu. 2020 · 2010
Cited alongside, same era.
High-dimensional structure estimation in ising models: Local separation criterion
A. Anandkumar, V. Y. F. Tan, F. Huang, and A. S. Willsky. 2012 · 2012
Cited alongside, same era.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and Jeffrey. 2013 · 2013
Cited alongside, same era.
Crowdsourcing a word-emotion association lexicon
S. M. Mohammad and P. D. Turney. 2013 · 2013
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
M. Peters, M. Neumann, L. Zettlemoyer, and W. Yih. 2018 · 2018
Later among the works it cites.
Analysis methods in neural language processing: A survey
Y. Belinkov and J. Glass. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova. 2019 · 2019
Later among the works it cites.
Syntactic dependencies correspond to word pairs with high mutual information
R. Futrell, P. Qian, E. Gibson, E. Fedorenko, and I. Blank. 2019 · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
J. Hewitt and C.D. Manning. 2019 · 2019
Later among the works it cites.
SemEval-2019 task 4: Hyperpartisan news detection
J. Kiesel, M. Mestre, R. Shukla, E. Vincent, P. Adineh, D. Corney, B. Stein, and M. Potthast. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Coordination structures in dependency treebanks
M. Popel, D. Mareček, J. Štěpánek, D. Zeman, and Z. Žabokrtský. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Y. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba. 2014 · 2014
Cited alongside, same era.
Neural word embedding as implicit matrix factorization
O. Levy and Y. Goldberg. 2014 · 2014
Cited alongside, same era.
The stanford coreNLP natural language processing toolkit
C. D. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J. Bethard, and D. McClosky. 2014 · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning. 2014 · 2014
Cited alongside, same era.
Random walks on context spaces: Towards an explanation of the mysteries of semantic word embeddings
S. Arora, Y. Li, Y. Liang, T. Ma, and A. Risteski. 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. 2019 · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
I. Tenney, P. Xia, B. Chen, A. Wang, A. Poliak, R. T. McCoy, N. Kim, B. Van Durme, S. Bowman, D. Das, and E. Pavlick. 2019 · 2019
Later among the works it cites.
Entity, relation, and event extraction with contextualized span representations
D. Wadden, U Wennberg, Y. Luan, and H. Hajishirzi. 2019 · 2019
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith. 2020 · 2020
Later among the works it cites.
How can we know what language models know?
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig. 2020 · 2020
Later among the works it cites.
Investigating transferability in pretrained language models
Alex Tamkin, Trisha Singh, Davide Giovanardi, and Noah Goodman. 2020 · 2020
Later among the works it cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Z. Wu, Y. Chen, B. Kao, and Q. Liu. 2020 · 2020
Later among the works it cites.
Incorporating BERT into neural machine translation
J. Zhu, Y. Xia, L. Wu, D. He, T. Qin, W. Zhou, H. Li, and T. Liu. 2020 · 2020
Later among the works it cites.