Fetching the paper…
Reading the bibliography…
In distributional semantics, the pointwise mutual information ($\mathit{PMI}$) weighting of the cooccurrence matrix performs far better than raw counts.
Distributional structure
Zellig S Harris. 1954 · 1954
Earlier work this paper cites.
A synopsis of linguistic theory, 1930-1955
John R Firth. 1957 · 1955
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Ward Church and Patrick Hanks. 1990 · 1990
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher D Manning, Christopher D Manning, and Hinrich Schütze. 1999 · 1999
Earlier work this paper cites.
Speech & language processing
Dan Jurafsky. 2000 · 2000
Earlier work this paper cites.
Improvements in automatic thesaurus extraction
James R. Curran and Marc Moens. 2002 · 2002
Earlier work this paper cites.
Measuring praise and criticism: Inference of semantic orientation from association
Peter D Turney and Michael L Littman. 2003 · 2003
Earlier work this paper cites.
Learning the unlearnable: The role of missing evidence
Terry Regier and Susanne Gahl. 2004 · 2004
Earlier work this paper cites.
Extracting semantic representations from word co-occurrence statistics: A computational study
John A Bullinaria and Joseph P Levy. 2007 · 2007
Cited alongside, same era.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma. 2009 · 2009
Cited alongside, same era.
Better word representations with recursive neural networks for morphology
Minh-Thang Luong, Richard Socher, and Christopher D Manning. 2013 · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
A systematic study of semantic vector space model parameters
Douwe Kiela and Stephen Clark. 2014 · 2014
Cited alongside, same era.
Neural word embedding as implicit matrix factorization
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Felix Hill, Roi Reichart, and Anna Korhonen. 2015 · 2015
Later among the works it cites.
Bidirectional lstm-crf models for sequence tagging
Zhiheng Huang, Wei Xu, and Kai Yu. 2015 · 2015
Later among the works it cites.
Robust co-occurrence quantification for lexical distributional semantics
Dmitrijs Milajevs, Mehrnoosh Sadrzadeh, and Matthew Purver. 2016 · 2016
Later among the works it cites.
Matrix factorization using window sampling and negative sampling for improved word representations
Alexandre Salle, Aline Villavicencio, and Marco Idiart. 2016 · 2016
Later among the works it cites.
Swivel: Improving embeddings by noticing what’s missing
Noam Shazeer, Ryan Doherty, Colin Evans, and Chris Waterson. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Omer Levy and Yoav Goldberg. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Cited alongside, same era.
Improving distributional semantic vectors through context selection and normalisation
Tamara Polajnar and Stephen Clark. 2014 · 2014
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Later among the works it cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Later among the works it cites.
Batch is not heavy: Learning word embeddings from all samples
Xin Xin, Yuan Fajie, and He Xiangnan. 2018 · 2018
Later among the works it cites.