Fetching the paper…
Reading the bibliography…
Softmax is an output activation function for modeling categorical probability distributions in many applications of deep learning.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Neural Networks for Pattern Recognition
Christopher M Bishop · 1995
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Numerical Recipes 3rd Edition: The Art of Scientific Computing
William H Press, Saul A Teukolsky, William T Vetterling, and Brian P Flannery · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Gated softmax classification
Roland Memisevic, Christopher Zach, Marc Pollefeys, and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Statistical language models based on neural networks
Tomas Mikolov · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Riemannian metrics for neural networks i: feedforward networks
Yann Ollivier · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo · 2016
Later among the works it cites.
One-vs-each approximation to softmax for scalable estimation of probabilities
Michalis K. Titsias · 2016
Later among the works it cites.
Noisy softmax: Improving the generalization ability of dcnn via postponing the early softmax saturation
Binghui Chen, Weihong Deng, and Junping Du · 2017
Later among the works it cites.
Efficient softmax approximation for GPUs
Édouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou · 2017
Later among the works it cites.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
An exploration of softmax alternatives belonging to the spherical loss family
Alexandre de Brébisson and Pascal Vincent · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters
John S Bridle
Cited in the paper.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
John S Bridle
Cited in the paper.
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Later among the works it cites.
Secureml: A system for scalable privacy-preserving machine learning
Payman Mohassel and Yupeng Zhang · 2017
Later among the works it cites.
SVD-softmax: Fast softmax approximation on large vocabulary neural networks
Kyuhong Shim, Minjae Lee, Iksoo Choi, Yoonho Boo, and Wonyong Sung · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Closest in time.
Breaking the softmax bottleneck: a high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen · 2018
Closest in time.