Fetching the paper…
Reading the bibliography…
Most unsupervised NLP models represent each word with a single point or single region in semantic space, while the existing multi-sense word embeddings cannot represent longer word sequences like phrases or sentences.
Multidimensional binary search trees used for associative searching
Bentley, J. L. 1975 · 1975
Earlier work this paper cites.
Distance measures for point sets and their computation
Eiter, T.; and Mannila, H. 1997 · 1997
Earlier work this paper cites.
Non-negative Sparse Coding
Hoyer, P. O. 2002 · 2002
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M.; Ng, A. Y.; and Jordan, M. I. 2003 · 2003
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Lin, C.-Y.; and Hovy, E. 2003 · 2003
Earlier work this paper cites.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2005
Earlier work this paper cites.
Sparse, Dense, and Attentional Representations for Text Retrieval
Luan, Y.; Eisenstein, J.; Toutanova, K.; and Collins, M. 2020 · 2005
Earlier work this paper cites.
Topic Models Conditioned on Arbitrary Features with Dirichlet-multinomial Regression
Mimno, D. M.; and McCallum, A. 2008 · 2008
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, S.; Klein, E.; and Loper, E. 2009 · 2009
Earlier work this paper cites.
Semeval-2012 task 6: A pilot on semantic textual similarity
Agirre, E.; Diab, M.; Cer, D.; and Gonzalez-Agirre, A. 2012 · 2012
Earlier work this paper cites.
Word sense induction for novel sense detection
Lau, J. H.; Cook, P.; McCarthy, D.; Newman, D.; and Baldwin, T. 2012 · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T.; and Hinton, G. 2012 · 2012
Earlier work this paper cites.
Domain and function: A dual-space model of semantic relations and compositions
Turney, P. D. 2012 · 2012
Earlier work this paper cites.
* SEM 2013 shared task: Semantic textual similarity
Agirre, E.; Cer, D.; Diab, M.; Gonzalez-Agirre, A.; and Guo, W. 2013 · 2013
Earlier work this paper cites.
Semeval-2013 task 5: Evaluating phrasal semantics
Korkontzelos, I.; Zesch, T.; Zanzotto, F. M.; and Biemann, C. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Semeval-2014 task 10: Multilingual semantic textual similarity
Agirre, E.; Banea, C.; Cardie, C.; Cer, D.; Diab, M.; Gonzalez-Agirre, A.; Guo, W.; Mihalcea, R.; Rigau, G.; and Wiebe, J. 2014 · 2014
Earlier work this paper cites.
Evaluating Neural Word Representations in Tensor-Based Compositional Settings
Milajevs, D.; Kartsaklis, D.; Sadrzadeh, M.; and Purver, M. 2014 · 2014
Earlier work this paper cites.
Efficient Non-parametric Estimation of Multiple Embeddings per Word in Vector Space
Neelakantan, A.; Shankar, J.; Passos, A.; and McCallum, A. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014 · 2014
Earlier work this paper cites.
Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability
Agirre, E.; Banea, C.; Cardie, C.; Cer, D.; Diab, M.; Gonzalez-Agirre, A.; Guo, W.; Lopez-Gazpio, I.; Maritxalar, M.; Mihalcea, R.; Rigau, G.; Uria, L.; and Wiebe, J. 2015 · 2015
Earlier work this paper cites.
A novel neural topic model and its supervised extension
Cao, Z.; Li, S.; Liu, Y.; Li, W.; and Ji, H. 2015 · 2015
Earlier work this paper cites.
Sparse Overcomplete Word Vector Representations
Faruqui, M.; Tsvetkov, Y.; Yogatama, D.; Dyer, C.; and Smith, N. A. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M.; Kocisky, T.; Grefenstette, E.; Espeholt, L.; Kay, W.; Suleyman, M.; and Blunsom, P. 2015 · 2015
Cited alongside, same era.
Skip-thought vectors
Kiros, R.; Zhu, Y.; Salakhutdinov, R. R.; Zemel, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Cited alongside, same era.
Summarization based on embedding distributions
Kobayashi, H.; Noguchi, M.; and Yatsuka, T. 2015 · 2015
Cited alongside, same era.
From word embeddings to document distances
Kusner, M.; Sun, Y.; Kolkin, N.; and Weinberger, K. 2015 · 2015
Cited alongside, same era.
PPDB 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification
Pavlick, E.; Rastogi, P.; Ganitkevitch, J.; Van Durme, B.; and Callison-Burch, C. 2015 · 2015
Cited alongside, same era.
Word Representations via Gaussian Embedding
Vilnis, L.; and McCallum, A. 2015 · 2015
Jointly Embedding Entities and Text with Distant Supervision
Newman-Griffis, D.; Lai, A. M.; and Fosler-Lussier, E. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Later among the works it cites.
Rezatofighi, S. H.; Kaskman, R.; Motlagh, F. T.; Shi, Q.; Cremers, D.; Leal-Taixé, L.; and Reid, I. 2018 · 2018
Later among the works it cites.
Compressing Word Embeddings via Deep Compositional Code Learning
Shu, R.; and Nakayama, H. 2018 · 2018
Later among the works it cites.
Loss Functions for Multiset Prediction
Welleck, S.; Yao, Z.; Gai, Y.; Mao, J.; Zhang, Z.; and Cho, K. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning composition models for phrase embeddings
Yu, M.; and Dredze, M. 2015 · 2015
Cited alongside, same era.
Semeval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation
Agirre, E.; Banea, C.; Cer, D.; Diab, M.; Gonzalez-Agirre, A.; Mihalcea, R.; Rigau, G.; and Wiebe, J. 2016 · 2016
Cited alongside, same era.
Ask the GRU: Multi-task Learning for Deep Text Recommendations
Bansal, T.; Belanger, D.; and McCallum, A. 2016 · 2016
Cited alongside, same era.
Improving Hypernymy Detection with an Integrated Path-based and Distributional Method
Shwartz, V.; Goldberg, Y.; and Dagan, I. 2016 · 2016
Cited alongside, same era.
End-to-end people detection in crowded scenes
Stewart, R.; Andriluka, M.; and Ng, A. Y. 2016 · 2016
Cited alongside, same era.
Multilingual Relation Extraction using Compositional Universal Schema
Verga, P.; Belanger, D.; Strubell, E.; Roth, B.; and McCallum, A. 2016 · 2016
Cited alongside, same era.
Asaadi, S.; Mohammad, S. M.; and Kiritchenko, S. 2019 · 2019
Later among the works it cites.
Holographic and other Point Set Distances for Machine Learning
Balles, L.; and Fischbacher, T. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
Insertion-based decoding with automatically inferred generation order
Gu, J.; Liu, Q.; and Cho, K. 2019 · 2019
Later among the works it cites.
Improving document classification with multi-sense embeddings
Gupta, V.; Saw, A.; Nokhiz, P.; Gupta, H.; and Talukdar, P. 2019 · 2019
Later among the works it cites.
Von Mises-Fisher Loss for Training Sequence to Sequence Models with Continuous Outputs
Kumar, S.; and Tsvetkov, Y. 2019 · 2019
Later among the works it cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Lee, J.; Lee, Y.; Kim, J.; Kosiorek, A. R.; Choi, S.; and Teh, Y. W. 2019 · 2019
Later among the works it cites.
Efficient Contextual Representation Learning With Continuous Outputs
Li, L. H.; Chen, P. H.; Hsieh, C.-J.; and Chang, K.-W. 2019 · 2019
Later among the works it cites.
L2g auto-encoder: Understanding point clouds by local-to-global reconstruction with hierarchical self-attention
Liu, X.; Han, Z.; Wen, X.; Liu, Y.-S.; and Zwicker, M. 2019 · 2019
Later among the works it cites.
Adapting RNN Sequence Prediction Model to Multi-label Set Prediction
Qin, K.; Li, C.; Pavlu, V.; and Aslam, J. A. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Later among the works it cites.
Insertion Transformer: Flexible Sequence Generation via Insertion Operations
Stern, M.; Chan, W.; Kiros, J.; and Uszkoreit, J. 2019 · 2019
Later among the works it cites.
Attention-based mixture density recurrent networks for history-based recommendation
Wang, T.; Cho, K.; and Wen, M. 2019 · 2019
Later among the works it cites.
Non-Monotonic Sequential Text Generation
Welleck, S.; Brantley, K.; Daumé III, H.; and Cho, K. 2019 · 2019
Later among the works it cites.
Sentence Centrality Revisited for Unsupervised Summarization
Zheng, H.; and Lapata, M. 2019 · 2019
Later among the works it cites.
P-SIF: Document embeddings using partition averaging
Gupta, V.; Saw, A.; Nokhiz, P.; Netrapalli, P.; Rai, P.; and Talukdar, P. 2020 · 2020
Later among the works it cites.
Context mover’s distance & barycenters: Optimal transport of contexts for building representations
Singh, S. P.; Hug, A.; Dieuleveut, A.; and Jaggi, M. 2020 · 2020
Later among the works it cites.
Changing the Mind of Transformers for Topically-Controllable Language Generation
Chang, H.-S.; Yuan, J.; Iyyer, M.; and McCallum, A. 2021 · 2021
Closest in time.
Multi-facet Universal Schema
Paul, R.; Chang, H.-S.; and McCallum, A. 2021 · 2021
Closest in time.