Fetching the paper…
Reading the bibliography…
Pre-trained word embeddings like ELMo and BERT contain rich syntactic and semantic information, resulting in state-of-the-art performance on various tasks.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aäron van den Oord, Alexander A. Alemi, and George Tucker. 2019 · 1905
Earlier work this paper cites.
Optimum Branchings
Jack Edmonds. 1966 · 1966
Earlier work this paper cites.
Distributional clustering of English words
Fernando Pereira, Naftali Tishby, and Lillian Lee. 1993 · 1993
Earlier work this paper cites.
Deterministic annealing for clustering, compression, classification, regression, and related optimization problems
Kenneth Rose. 1998 · 1998
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek. 2000 · 2000
Earlier work this paper cites.
Multivariate information bottleneck
Nir Friedman, Ori Mosenzon, Noam Slonim, and Naftali Tishby. 2001 · 2001
Earlier work this paper cites.
An Introduction to Multivariate Statistical Analysis
T.W. Anderson. 2003 · 2003
Earlier work this paper cites.
Accurate unlexicalized parsing
D. Klein and C. D. Manning. 2003 · 2003
Earlier work this paper cites.
Structured prediction models via the matrix-tree theorem
Terry Koo, Amir Globerson, Xavier Carreras Pérez, and Michael Collins. 2007 · 2007
Earlier work this paper cites.
On the complexity of non-projective data-driven dependency parsing
Ryan McDonald and Giorgio Satta. 2007 · 2007
Earlier work this paper cites.
Probabilistic models of nonprojective dependency trees
David A Smith and Noah A Smith. 2007 · 2007
Cited alongside, same era.
Visualizing high-dimensional data using t-SNE
L.J.P. van der Maaten and G.E. Hinton. 2008 · 2008
Cited alongside, same era.
Natural Language Processing with Python , 1st edition
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2016 · 2016
Later among the works it cites.
AllenNLP: A Deep Semantic Natural Language Processing Platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke S. Zettlemoyer. 2017 · 2017
Later among the works it cites.
Towards better UD parsing: Deep contextualized word embeddings, ensemble, and treebank concatenation
Wanxiang Che, Yijia Liu, Yuxuan Wang, Bo Zheng, and Ting Liu. 2018 · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. 2014 · 2014
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015b · 2015
Cited alongside, same era.
Deep Variational Information Bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, and Kevin Murphy. 2016 · 2016
Cited alongside, same era.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D. Manning. 2016 · 2016
Cited alongside, same era.
Categorical reparameterization with Gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018b
Cited in the paper.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015a
Cited in the paper.
Later among the works it cites.
Universal dependencies 2.3
Joakim Nivre et al. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018a · 2018
Later among the works it cites.
Efficient human-like semantic representations via the information bottleneck principle
Noga Zaslavsky, Charles Kemp, Terry Regier, and Naftali Tishby. 2018 · 2018
Later among the works it cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Closest in time.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Closest in time.