Fetching the paper…
Reading the bibliography…
We study the problem of learning similarity functions over very large corpora using neural network embedding models.
Monte Carlo Methods
J. Hammersley and D. Handscomb · 1964
Earlier work this paper cites.
Signature verification using a "siamese" time delay neural network
J. Bromley, J. W. Bentz, L. Bottou, I. Guyon, Y. LeCun, C. Moore, E. Säckinger, and R. Shah · 1993
Earlier work this paper cites.
Approximating Integrals via Monte Carlo and Deterministic Methods
M. Evans and T. Swartz · 2000
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Y. Bengio and J. Senecal · 2003
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Y. Bengio and J. Senecal · 2008
Earlier work this paper cites.
Collaborative filtering for implicit feedback datasets
Y. Hu, Y. Koren, and C. Volinsky · 2008
Earlier work this paper cites.
Regression-based latent factor models
D. Agarwal and B.-C. Chen · 2009
Earlier work this paper cites.
Large scale online learning of image similarity through ranking
G. Chechik, V. Sharma, U. Shalit, and S. Bengio · 2010
Earlier work this paper cites.
Factorization machines
S. Rendle · 2010
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
Neural word embedding as implicit matrix factorization
O. Levy and Y. Goldberg · 2014
Cited alongside, same era.
Smoothed gradients for stochastic variational inference
S. Mandt and D. Blei · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Cited alongside, same era.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
A. Shrivastava and P. Li · 2014
Cited alongside, same era.
The movielens datasets: History and context
F. M. Harper and J. A. Konstan · 2015
Cited alongside, same era.
An exploration of softmax alternatives belonging to the spherical loss family
A. de Brébisson and P. Vincent · 2016
Later among the works it cites.
Node2vec: Scalable feature learning for networks
A. Grover and J. Leskovec · 2016
Later among the works it cites.
Swivel: Improving embeddings by noticing what’s missing
N. Shazeer, R. Doherty, C. Evans, and C. Waterson · 2016
Later among the works it cites.
Tapas: Two-pass approximate adaptive sampling for softmax
Y. Bai, S. Goldman, and L. Zhang · 2017
Later among the works it cites.
A generic coordinate descent framework for learning from implicit feedback
I. Bayer, X. He, B. Kanagal, and S. Rendle · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On symmetric and asymmetric lshs for inner product search
B. Neyshabur and N. Srebro · 2015
Cited alongside, same era.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Cited alongside, same era.
Efficient exact gradient update for training deep networks with very large sparse targets
P. Vincent, A. de Brébisson, and X. Bouthillier · 2015
Cited alongside, same era.
Strategies for training large vocabulary neural language models
W. Chen, D. Grangier, and M. Auli · 2016
Cited alongside, same era.
Wikimedia downloads
Wikimedia Foundation
Cited in the paper.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Later among the works it cites.
Folding: Why good models sometimes make spurious recommendations
D. Xin, N. Mayoraz, H. Pham, K. Lakshmanan, and J. R. Anderson · 2017
Later among the works it cites.
Selection of negative samples for one-class matrix factorization
H.-F. Yu, M. Bilenko, and C.-J. Lin · 2017
Later among the works it cites.
Network embedding as matrix factorization: Unifying deepwalk, line, pte, and node2vec
J. Qiu, Y. Dong, H. Ma, J. Li, K. Wang, and J. Tang · 2018
Closest in time.