Fetching the paper…
Reading the bibliography…
In a multi-class classification problem, it is standard to model the output of a neural network as a categorical distribution conditioned on the inputs.
Building a large annotated corpus of english: The penn treebank
Marcus, Mitchell P, Marcinkiewicz, Mary Ann, and Santorini, Beatrice · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Yoshua, Ducharme, Réjean, and Vincent, Pascal · 2001
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, Frederic and Bengio, Yoshua · 2005
Earlier work this paper cites.
Efficient bounds for the softmax function and applications to approximate inference in hybrid models
Bouchard, Guillaume · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex and Hinton, Geoffrey · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvarinen, A · 2010
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., and Dean, J · 2013
Cited alongside, same era.
Learning word embeddings efficiently with noise-contrastive estimation
Mnih, Andriy and Kavukcuoglu, Koray · 2013
Cited alongside, same era.
Riemannian metrics for neural networks
Ollivier, Yann · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Sutskever, Ilya, Martens, James, Dahl, George, and Hinton, Geoffrey · 2013
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Hill, Felix, Reichart, Roi, and Korhonen, Anna · 2014
Later among the works it cites.
Asymmetric LSH (ALSH) for sublinear time maximum inner product search (MIPS)
Shrivastava, Anshumali and Li, Ping · 2014
Later among the works it cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chelba, Ciprian, Mikolov, Tomas, Schuster, Mike, Ge, Qi, Brants, Thorsten, Koehn, Phillipp, and Robinson, Tony · 2014
Cited alongside, same era.
Efficient exact gradient update for training deep networks with very large sparse targets
Vincent, Pascal, de Brébisson, Alexandre, and Bouthillier, Xavier · 2015
Closest in time.