Fetching the paper…
Reading the bibliography…
We present a framework for building unsupervised representations of entities and their compositions, where each entity is viewed as a probability distribution rather than a vector embedding.
On the translocation of masses
Leonid V Kantorovich · 1942
Earlier work this paper cites.
Distributional structure
Zellig S Harris · 1954
Earlier work this paper cites.
A relationship between arbitrary positive matrices and doubly stochastic matrices
Richard Sinkhorn · 1964
Earlier work this paper cites.
Contextual correlates of synonymy
Herbert Rubenstein and John B Goodenough · 1965
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Ward Church and Patrick Hanks · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
A general framework for distributional similarity
Julie Weeds and David Weir · 2003
Earlier work this paper cites.
The distributional inclusion hypotheses and lexical entailment
Maayan Geffet and Ido Dagan · 2005
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Vector-based models of semantic composition
Jeff Mitchell and Mirella Lapata · 2008
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Cédric Villani · 2008
Earlier work this paper cites.
The wacky wide web: a collection of very large linguistically processed web-crawled corpora
Marco Baroni, Silvia Bernardini, Adriano Ferraresi, and Eros Zanchetta · 2009
Earlier work this paper cites.
Directional distributional similarity for lexical inference
Lili Kotlerman, Ido Dagan, Idan Szpektor, and Maayan Zhitomirsky-Geffet · 2010
Earlier work this paper cites.
Barycenters in the wasserstein space
Martial Agueh and Guillaume Carlier · 2011
Earlier work this paper cites.
How we blessed distributional semantic evaluation
Marco Baroni and Alessandro Lenci · 2011
Earlier work this paper cites.
Semeval-2012 task 6: A pilot on semantic textual similarity
Eneko Agirre, Mona Diab, Daniel Cer, and Aitor Gonzalez-Agirre · 2012
Earlier work this paper cites.
Improving word representations via global context and multiple word prototypes
Eric H Huang, Richard Socher, Christopher D Manning, and Andrew Y Ng · 2012
Earlier work this paper cites.
* sem 2013 shared task: Semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo · 2013
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Semeval-2014 task 10: Multilingual semantic textual similarity
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe · 2014
Earlier work this paper cites.
Fast computation of wasserstein barycenters
Marco Cuturi and Arnaud Doucet · 2014
Earlier work this paper cites.
Community evaluation and exchange of word vectors at wordvectors.org
Manaal Faruqui and Chris Dyer · 2014
Earlier work this paper cites.
Learning sense-specific word embeddings by exploiting bilingual resources
Jiang Guo, Wanxiang Che, Haifeng Wang, and Ting Liu · 2014
Earlier work this paper cites.
A Convolutional Neural Network for Modelling Sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom · 2014
Earlier work this paper cites.
Convolutional Neural Networks for Sentence Classification
Yoon Kim · 2014
Cited alongside, same era.
The stanford corenlp natural language processing toolkit
Christopher Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven Bethard, and David McClosky · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
Distributional lexical entailment by topic coherence
Laura Rimell · 2014
Cited alongside, same era.
Chasing hypernyms in vector spaces with entropy
Enrico Santus, Alessandro Lenci, Qin Lu, and S Schulte im Walde · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Cited alongside, same era.
Improving hypernymy detection with an integrated path-based and distributional method
Vered Shwartz, Yoav Goldberg, and Ido Dagan · 2016
Later among the works it cites.
Near-linear time approximation algorithms for optimal transport via sinkhorn iteration
Jason Altschuler, Jonathan Weed, and Philippe Rigollet · 2017
Later among the works it cites.
A simple but tough-to-beat baseline for sentence embeddings
Sanjeev Arora, Yingyu Liang, and Tengyu Ma · 2017
Later among the works it cites.
Ben Athiwaratkun and Andrew Gordon Wilson · 2017
Later among the works it cites.
Distributional inclusion vector embedding for unsupervised hypernymy detection
Haw-Shiuan Chang, ZiYun Wang, Luke Vilnis, and Andrew McCallum · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luke Vilnis and Andrew McCallum · 2014
Cited alongside, same era.
Learning to distinguish hypernyms and co-hyponyms
Julie Weeds, Daoud Clarke, Jeremy Reffin, David Weir, and Bill Keller · 2014
Cited alongside, same era.
Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, et al · 2015
Cited alongside, same era.
Iterative bregman projections for regularized transportation problems
Jean-David Benamou, Guillaume Carlier, Marco Cuturi, Luca Nenna, and Gabriel Peyré · 2015
Cited alongside, same era.
Distributional models for semantic relations: A study on hyponymy and antonymy
Giulia Benotto · 2015
Cited alongside, same era.
E-commerce in your inbox: Product recommendations at scale
Mihajlo Grbovic, Vladan Radosavljevic, Nemanja Djuric, Narayan Bhamidipati, Jaikit Savla, Varun Bhagwan, and Doug Sharp · 2015
Cited alongside, same era.
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes · 2017
Later among the works it cites.
Leveraging Large Amounts of Weakly Supervised Data for Multi-Language Sentiment Classification
Jan Deriu, Aurelien Lucchi, Valeria De Luca, Aliaksei Severyn, Simon Müller, Mark Cieliebak, Thomas Hofmann, and Martin Jaggi · 2017
Later among the works it cites.
Learning word embeddings for hyponymy with entailment-based distributional semantics
James Henderson · 2017
Later among the works it cites.
Unsupervised learning of sentence embeddings using compositional n-gram features
Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi · 2017
Later among the works it cites.
Hyperlex: A large-scale evaluation of graded lexical entailment
Ivan Vulić, Daniela Gerz, Douwe Kiela, Felix Hill, and Anna Korhonen · 2017
Later among the works it cites.
Starspace: Embed all the things!
Ledell Wu, Adam Fisch, Sumit Chopra, Keith Adams, Antoine Bordes, and Jason Weston · 2017
Later among the works it cites.
Determining gains acquired from word embedding quantitatively using discrete distribution clustering
Jianbo Ye, Yanran Li, Zhaohui Wu, James Z Wang, Wenjie Li, and Jia Li · 2017
Later among the works it cites.
Earth mover’s distance minimization for unsupervised bilingual lexicon induction
Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun · 2017
Later among the works it cites.
Findings of the 2018 conference on machine translation (WMT18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz · 2018
Closest in time.
Universal sentence encoder, 2018
Daniel Cer, Yinfei Yang, Sheng yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil · 2018
Closest in time.
SentEval: An Evaluation Toolkit for Universal Sentence Representations
Alexis Conneau and Douwe Kiela · 2018
Closest in time.
Unsupervised Alignment of Embeddings with Wasserstein Procrustes
Edouard Grave, Armand Joulin, and Quentin Berthet · 2018
Closest in time.
Inferlite: Simple universal sentence representations from natural language inference data
Jamie Kiros and William Chan · 2018
Closest in time.
Generalizing Point Embeddings using the Wasserstein Space of Elliptical Distributions
Boris Muzellec and Marco Cuturi · 2018
Closest in time.
Learning general purpose distributed sentence representations via large scale multi-task learning, 2018
Sandeep Subramanian, Adam Trischler, Yoshua Bengio, and Christopher J Pal · 2018
Closest in time.
Gaussian Word Embedding with a Wasserstein Distance Loss
Chi Sun, Hang Yan, Xipeng Qiu, and Xuanjing Huang · 2018
Closest in time.
Poincar \ \backslash ’e glove: Hyperbolic word embeddings
Alexandru Tifrea, Gary Bécigneul, and Octavian-Eugen Ganea · 2018
Closest in time.
Word mover’s embedding: From word2vec to document embedding
Lingfei Wu, Ian En-Hsu Yen, Kun Xu, Fangli Xu, Avinash Balakrishnan, Pin-Yu Chen, Pradeep Ravikumar, and Michael J. Witbrock · 2018
Closest in time.
Distilled wasserstein learning for word embedding and topic modeling
Hongteng Xu, Wenlin Wang, Wei Liu, and Lawrence Carin · 2018
Closest in time.
Universal Sentence Encoder results on STS 12-16
BERT official repo · 2019
Closest in time.
Learning embeddings into entropic wasserstein spaces, 2019
Charlie Frogner, Farzaneh Mirzazadeh, and Justin Solomon · 2019
Closest in time.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig · 2084
Closest in time.