Fetching the paper…
Reading the bibliography…
Softmax is the most commonly used output function for multiclass problems and is widely used in areas such as vision, natural language processing, and recommendation.
An efficient method for generating discrete random variables with general distributions
Walker, A. J · 1977
Earlier work this paper cites.
Treebank-3 ldc99t42, 1999
Marcus, M., Santorini, B., Marcinkiewicz, M. A., and Taylor, A · 1999
Earlier work this paper cites.
Classes for fast maximum entropy training
Goodman, J · 2001
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Bengio, Y. and Sénécal, J.-S · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, F. and Bengio, Y · 2005
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Bengio, Y. and Sénécal, J.-S · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Structured output layer neural network language model
Le, H.-S., Oparin, I., Allauzen, A., Gauvain, J.-L., and Yvon, F · 2011
Cited alongside, same era.
Extensions of recurrent neural network language model
Mikolov, T., Kombrink, S., Burget, L., Cernocký, J., and Khudanpur, S · 2011
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G., and Dean, J · 2013
Cited alongside, same era.
Speed regularization and optimality in word classing
Zweig, G. and Makarychev, K · 2013
Cited alongside, same era.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Shrivastava, A. and Li, P · 2014
Cited alongside, same era.
An exploration of softmax alternatives belonging to the spherical loss family
Brébisson, A. d. and Vincent, P · 2015
Later among the works it cites.
Strategies for training large vocabulary neural language models
Chen, W., Grangier, D., and Auli, M · 2015
Later among the works it cites.
Sublinear partition estimation
Rastogi, P. and Durme, B. V · 2015
Later among the works it cites.
Efficient exact gradient update for training deep networks with very large sparse targets
Vincent, P., Brébisson, A. d., and Bouthillier, X · 2015
Later among the works it cites.
Deep neural networks for youtube recommendations
Covington, P., Adams, J., and Sargin, E · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Cited alongside, same era.
Clustering is efficient for approximate maximum inner product search
Auvolat, A. and Vincent, P · 2015
Cited alongside, same era.
Speeding up neural networks for large scale classification using WTA hashing
Bakhtiary, A. H., Lapedriza, À., and Masip, D · 2015
Cited alongside, same era.
TAPAS: two-pass approximate adaptive sampling for softmax
Bai, Y., Goldman, S., and Zhang, L · 2017
Closest in time.
Efficient softmax approximation for GPUs
Grave, É., Joulin, A., Cissé, M., Grangier, D., and Jégou, H · 2017
Closest in time.
An experimental analysis of noise-contrastive estimation: the noise distribution matters
Labeau, M. and Allauzen, A · 2017
Closest in time.