Fetching the paper…
Reading the bibliography…
Distributed representations of words encode lexical semantic information, but what type of information is encoded and how? Focusing on the skip-gram with negative-sampling method, we found that the squared norm of static word embedding encodes the information gain conveyed by the word; the information gain is defined by the Kullback-Leibler divergence of the co-occurrence distribution of the word to the unigram distribution.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
A synopsis of linguistic theory 1930-55
J. R. Firth. 1957 · 1952
Earlier work this paper cites.
Distributional structure
Zellig Harris. 1954 · 1954
Earlier work this paper cites.
The geometry of exponential families
Bradley Efron. 1978 · 1978
Earlier work this paper cites.
Differential Geometry of Curved Exponential Families-Curvatures and Information Loss
Shun-Ichi Amari. 1982 · 1982
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari. 1998 · 1998
Earlier work this paper cites.
Theory of point estimation
Erich L Lehmann and George Casella. 1998 · 1998
Earlier work this paper cites.
The NLM indexing initiative
A. R. Aronson, O. Bodenreider, H. F. Chang, S. M. Humphrey, J. G. Mork, S. J. Nelson, T. C. Rindflesch, and W. J. Wilbur. 2000 · 2000
Earlier work this paper cites.
Improved automatic keyword extraction given more linguistic knowledge
Anette Hulth. 2003 · 2003
Earlier work this paper cites.
A general framework for distributional similarity
Julie Weeds and David Weir. 2003 · 2003
Earlier work this paper cites.
Keyword extraction from a single document using word co-occurrence statistical information
Y. Matsuo and M. Ishizuka. 2004 · 2004
Earlier work this paper cites.
Characterising measures of lexical distributional similarity
Julie Weeds, David Weir, and Diana McCarthy. 2004 · 2004
Earlier work this paper cites.
The statistics of word cooccurrences: word pairs and collocations
Stefan Evert. 2005 · 2005
Earlier work this paper cites.
Keyphrase extraction in scientific publications
Thuy Dung Nguyen and Min-Yen Kan. 2007 · 2007
Earlier work this paper cites.
Domain-independent automatic keyphrase indexing with small training sets
Olena Medelyan and Ian H. Witten. 2008 · 2008
Earlier work this paper cites.
Topic indexing with Wikipedia
Olena Medelyan, Ian H Witten, and David Milne. 2008 · 2008
Earlier work this paper cites.
Keyphrase extraction from single documents in the open domain exploiting linguistic and statistical methods
A. T. Schutz. 2008 · 2008
Earlier work this paper cites.
PCA consistency in high dimension, low sample size context
Sungkyu Jung and J Stephen Marron. 2009 · 2009
Earlier work this paper cites.
Large dataset for keyphrases extraction
Mikalai Krapivin, Aliaksandr Autaeu, and Maurizio Marchese. 2009 · 2009
Earlier work this paper cites.
Human-competitive tagging using automatic keyphrase extraction
Olena Medelyan, Eibe Frank, and Ian H. Witten. 2009 · 2009
Earlier work this paper cites.
SemEval-2010 task 5 : Automatic keyphrase extraction from scientific articles
Su Nam Kim, Olena Medelyan, Min-Yen Kan, and Timothy Baldwin. 2010 · 2010
Earlier work this paper cites.
Composition in distributional models of semantics
Jeff Mitchell and Mirella Lapata. 2010 · 2010
Earlier work this paper cites.
Towards the quantification of the semantic information encoded in written language
Marcelo A. Montemurro and Damiá n H. Zanette. 2010 · 2010
Cited alongside, same era.
Keyword extraction using word co-occurrence
Christian Wartena, Rogier Brussee, and Wout Slakhorst. 2010 · 2010
Cited alongside, same era.
How we BLESSed distributional semantic evaluation
Marco Baroni and Alessandro Lenci. 2011 · 2011
Cited alongside, same era.
About the test data
Matt Mahoney. 2011 · 2011
Cited alongside, same era.
Keyphrase cloud generation of broadcast news
Luís Marujo, Márcio Viveiros, and João Paulo da Silva Neto. 2011 · 2011
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Michael Gutmann and Aapo Hyvärinen. 2012 · 2012
Cited alongside, same era.
SemEval 2017 task 10: ScienceIE - extracting keyphrases and relations from scientific publications
Isabelle Augenstein, Mrinal Das, Sebastian Riedel, Lakshmi Vikraman, and Andrew McCallum. 2017 · 2017
Later among the works it cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Later among the works it cites.
Investigating different syntactic context types and context representations for learning word embeddings
Bofang Li, Tao Liu, Zhe Zhao, Buzhou Tang, Aleksandr Drozd, Anna Rogers, and Xiaoyong Du. 2017 · 2017
Later among the works it cites.
Hypernyms under siege: Linguistically-motivated artillery for hypernymy detection
Vered Shwartz, Enrico Santus, and Dominik Schlechtweg. 2017 · 2017
Later among the works it cites.
A survey of high dimension low sample size asymptotics
Makoto Aoshima, Dan Shen, Haipeng Shen, Kazuyoshi Yata, Yi-Hui Zhou, and James S Marron. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Identifying hypernyms in distributional semantic spaces
Alessandro Lenci and Giulia Benotto. 2012 · 2012
Cited alongside, same era.
Categorical Data Analysis , 3rd edition
Alan Agresti. 2013 · 2013
Cited alongside, same era.
Measuring semantic content in distributional vectors
Aurélie Herbelot and Mohan Ganesalingam. 2013 · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomás Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
Information and exponential families: in statistical theory
Ole Barndorff-Nielsen. 2014 · 2014
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomás Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2014 · 2014
Cited alongside, same era.
Nikolay Arefyev, Pavel Ermolaev, and Alexander Panchenko. 2018 · 2018
Later among the works it cites.
A la carte embedding: Cheap but effective induction of semantic feature vectors
Mikhail Khodak, Nikunj Saunshi, Yingyu Liang, Tengyu Ma, Brandon Stewart, and Sanjeev Arora. 2018 · 2018
Later among the works it cites.
Unsupervised learning of sentence embeddings using compositional n-gram features
Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi. 2018 · 2018
Later among the works it cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
The connection between bayesian inference and information theory for model selection, information gain and experimental design
Sergey Oladyshkin and Wolfgang Nowak. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Later among the works it cites.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Word rotator’s distance
Sho Yokoi, Ryo Takahashi, Reina Akama, Jun Suzuki, and Kentaro Inui. 2020 · 2020
Later among the works it cites.
More than just frequency? demasking unsupervised hypernymy prediction methods
Thomas Bott, Dominik Schlechtweg, and Sabine Schulte im Walde. 2021 · 2021
Later among the works it cites.
Statistical Universals of Language
Kumiko Tanaka-Ishii. 2021 · 2021
Later among the works it cites.
English wikipedia dump data
Wikimedia Foundation. 2021 · 2021
Later among the works it cites.
What company do words keep? revisiting the distributional semantics of J.R. firth & zellig Harris
Mikael Brunila and Jack LaViolette. 2022 · 2022
Closest in time.
Exponential Families in Theory and Practice
Bradley Efron. 2022 · 2022
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.