Fetching the paper…
Reading the bibliography…
We propose a novel Wasserstein method with a distillation mechanism, yielding joint learning of word embeddings and topics.
Introduction to modern information retrieval
S. Gerard and J. M. Michael · 1983
Earlier work this paper cites.
Learning to classify text using support vector machines: Methods, theory and algorithms
T. Joachims · 2002
Earlier work this paper cites.
Laplacian Eigenmaps for dimensionality reduction and data representation
M. Belkin and P. Niyogi · 2003
Earlier work this paper cites.
Latent Dirichlet allocation
D. M. Blei, A. Y. Ng, and M. I. Jordan · 2003
Earlier work this paper cites.
Optimal transport: Old and new
C. Villani · 2008
Earlier work this paper cites.
Barycenters in the Wasserstein space
M. Agueh and G. Carlier · 2011
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
M. Cuturi · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Fast computation of Wasserstein barycenters
M. Cuturi and A. Doucet · 2014
Earlier work this paper cites.
Distributed representations of sentences and documents
Q. Le and T. Mikolov · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Iterative Bregman projections for regularized transportation problems
J.-D. Benamou, G. Carlier, M. Cuturi, L. Nenna, and G. Peyré · 2015
Earlier work this paper cites.
Distribution’s template estimate with Wasserstein metrics
E. Boissard, T. Le Gouic, J.-M. Loubes, et al · 2015
Earlier work this paper cites.
Distilling knowledge from deep networks with applications to healthcare domain
Z. Che, S. Purushotham, R. Khemani, and Y. Liu · 2015
Earlier work this paper cites.
Gaussian LDA for topic models with word embeddings
R. Das, M. Zaheer, and C. Dyer · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Cited alongside, same era.
From word embeddings to document distances
M. Kusner, Y. Sun, N. Kolkin, and K. Weinberger · 2015
Cited alongside, same era.
Topical word embeddings
Y. Liu, Z. Liu, T.-S. Chua, and M. Sun · 2015
Cited alongside, same era.
Unifying distillation and privileged information
D. Lopez-Paz, L. Bottou, B. Schölkopf, and V. Vapnik · 2015
Cited alongside, same era.
Multi-layer representation learning for medical concepts
E. Choi, M. T. Bahadori, E. Searles, C. Coffey, M. Thompson, J. Bost, J. Tejedor-Sojo, and J. Sun · 2016
Cited alongside, same era.
Cross modal distillation for supervision transfer
Sinkhorn-AutoDiff: Tractable Wasserstein learning of generative models
A. Genevay, G. Peyré, and M. Cuturi · 2017
Later among the works it cites.
Multitask learning and benchmarking with clinical time series data
H. Harutyunyan, H. Khachatrian, D. C. Kale, and A. Galstyan · 2017
Later among the works it cites.
Linear ensembles of word embedding models
A. Muromägi, K. Sirts, and S. Laur · 2017
Later among the works it cites.
Regularizing neural networks by penalizing confident output distributions
G. Pereyra, G. Tucker, J. Chorowski, Ł. Kaiser, and G. Hinton · 2017
Later among the works it cites.
Jointly learning word embeddings and latent topics
B. Shi, W. Lam, S. Jameel, S. Schockaert, and K. P. Lai · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Gupta, J. Hoffman, and J. Malik · 2016
Cited alongside, same era.
Supervised word mover’s distance
G. Huang, C. Guo, M. J. Kusner, Y. Sun, F. Sha, and K. Q. Weinberger · 2016
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
H. Inan, K. Khosravi, and R. Socher · 2016
Cited alongside, same era.
MIMIC-III, a freely accessible critical care database
A. E. Johnson, T. J. Pollard, L. Shen, H. L. Li-wei, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark · 2016
Cited alongside, same era.
A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell · 2016
Cited alongside, same era.
Learning to learn: Model regression networks for easy small sample learning
Y.-X. Wang and M. Hebert · 2016
Cited alongside, same era.
Near-linear time approximation algorithms for optimal transport via Sinkhorn iteration
J. Altschuler, J. Weed, and P. Rigollet · 2017
Cited alongside, same era.
Later among the works it cites.
Towards automated ICD coding using deep learning
H. Shi, P. Xie, Z. Hu, M. Zhang, and E. P. Xing · 2017
Later among the works it cites.
Topic compositional neural language model
W. Wang, Z. Gan, W. Wang, D. Shen, J. Huang, W. Ping, S. Satheesh, and L. Carin · 2017
Later among the works it cites.
Fast discrete distribution clustering using Wasserstein barycenter with sparse support
J. Ye, P. Wu, J. Z. Wang, and J. Li · 2017
Later among the works it cites.
Fréchet means and Procrustes analysis in Wasserstein space
Y. Zemel and V. M. Panaretos · 2017
Later among the works it cites.
J. M. Bajor, D. A. Mesa, T. J. Osterman, and T. A. Lasko · 2018
Closest in time.
An empirical evaluation of deep learning for ICD-9 code assignment using MIMIC-III clinical notes
J. Huang, C. Osorio, and L. W. Sy · 2018
Closest in time.
Explainable prediction of medical codes from clinical text
J. Mullenbach, S. Wiegreffe, J. Duke, J. Sun, and J. Eisenstein · 2018
Closest in time.
Wasserstein dictionary learning: Optimal transport-based unsupervised nonlinear dictionary learning
M. A. Schmitz, M. Heitz, N. Bonneel, F. Ngole, D. Coeurjolly, M. Cuturi, G. Peyré, and J.-L. Starck · 2018
Closest in time.
Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms
D. Shen, G. Wang, W. Wang, M. R. Min, Q. Su, Y. Zhang, C. Li, R. Henao, and L. Carin · 2018
Closest in time.