Fetching the paper…
Reading the bibliography…
The output scores of a neural network classifier are converted to probabilities via normalizing over the scores of all competing categories.
Refinements to nearest-neighbor searching in k-dimensional trees
R. F. Sproull · 1991
Earlier work this paper cites.
Neural networks for classification: a survey
G. P. Zhang · 2000
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond (Adaptive Computation and Machine Learning)
B. Schölkopf and A. J. Smola · 2001
Earlier work this paper cites.
Regularization with dot-product kernels
A. J. Smola, Z. L. Óvári, and R. C. Williamson · 2001
Earlier work this paper cites.
Asymptotics of sums of lognormal random variables with gaussian copula
S. Asmussen and L. Rojas-Nandayapa · 2008
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Y. Bengio and J.-S. Sénécal · 2008
Earlier work this paper cites.
Modeling lsh for performance tuning
W. Dong, Z. Wang, W. Josephson, M. Charikar, and K. Li · 2008
Earlier work this paper cites.
A scalable hierarchical distributed language model
A. Mnih and G. E. Hinton · 2008
Earlier work this paper cites.
Fast approximate nearest neighbors with automatic algorithm configuration
M. Muja and D. G. Lowe · 2009
Earlier work this paper cites.
What does classifying more than 10,000 image categories tell us?
J. Deng, A. C. Berg, K. Li, and L. Fei-Fei · 2010
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvärinen · 2010
Cited alongside, same era.
On the difficulty of nearest neighbor search
J. He, S. Kumar, and S.-f. Chang · 2012
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Cited alongside, same era.
Random feature maps for dot product kernels
P. Kar and H. Karnick · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
A. Mnih and Y. W. Teh · 2012
Speeding up the xbox recommender system using a euclidean transformation for inner-product spaces
Y. Bachrach, Y. Finkelstein, R. Gilad-Bachrach, L. Katzir, N. Koenigstein, N. Nice, and U. Paquet · 2014
Later among the works it cites.
Fast and robust neural network joint models for statistical machine translation
J. Devlin, R. Zbib, Z. Huang, T. Lamar, R. Schwartz, and J. Makhoul · 2014
Later among the works it cites.
Scalable nearest neighbour algorithms for high dimensional data
M. Muja and D. Lowe · 2014
Later among the works it cites.
On Symmetric and Asymmetric LSHs for Inner Product Search
B. Neyshabur and N. Srebro · 2014
Later among the works it cites.
ImageNet: Large Scale Visual Recognition Challenge, 2014
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Later among the works it cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Fast and scalable polynomial kernels via explicit feature maps
N. Pham and R. Pagh · 2013
Cited alongside, same era.
A. Shrivastava and P. Li · 2014
Later among the works it cites.
Asymmetric Minwise Hashing
A. Shrivastava and P. Li · 2014
Later among the works it cites.
When and why are log-linear models self-normalizing?
J. Andreas and D. Klein · 2015
Closest in time.