Fetching the paper…
Reading the bibliography…
Despite being the standard loss function to train multi-class neural networks, the log-softmax has two potential limitations.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, and P. Vincent · 2001
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
F. Morin and Y. Bengio · 2005
Earlier work this paper cites.
Ranking with ordered weighted pairwise classification
D. B. Nicolas Usunier and P. Gallinari · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvarinen · 2010
Earlier work this paper cites.
Wsabie: Scaling up to large vocabulary image annotation
J. Weston, S. Bengio, and N. Usunier · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Learning word embeddings efficiently with noise-contrastive estimation
A. Mnih and K. Kavukcuoglu · 2013
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2014
Cited alongside, same era.
Asymmetric LSH (ALSH) for sublinear time maximum inner product search (MIPS)
A. Shrivastava and P. Li · 2014
Cited alongside, same era.
Efficient exact gradient update for training deep networks with very large sparse targets
P. Vincent, A. d. Brébisson, and X. Bouthillier · 2015
Cited alongside, same era.
Top-k multiclass svm
M. H. Maksim Lapin and B. Schiele · 2015
Later among the works it cites.
Strategies for training large vocabulary neural language models
W. Chen, D. Grangier, and M. Auli · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Later among the works it cites.
An exploration of softmax alternatives belonging to the spherical loss family
A. d. Brébisson and P. Vincent · 2016
Closest in time.
Blackout: Speeding up recurrent neural network language models with very large vocabularies
S. Ji, S. Vishwanathan, N. Satish, M. J. Anderson, and P. Dubey · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…