Fetching the paper…
Reading the bibliography…
The softmax representation of probabilities for categorical variables plays a prominent role in modern machine learning with numerous applications in areas such as large scale classification, neural language modeling and recommendation systems.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Bradley, R. A. and Terry, M. E. (1952) · 1952
Earlier work this paper cites.
Multinomial logistic regression algorithm
Bohning, D. (1992) · 1992
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Bengio, Y. and Sénécal, J.-S. (2003) · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, F. and Bengio, Y. (2005) · 2005
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Bishop, C. M. (2006) · 2006
Earlier work this paper cites.
Generalized Bradley-Terry models and multi-class probability estimates
Huang, T.-K., Weng, R. C., and Lin, C.-J. (2006) · 2006
Earlier work this paper cites.
Efficient bounds for the softmax function and applications to approximate inference in hybrid models
Bouchard, G. (2007) · 2007
Earlier work this paper cites.
Multilabel text classification for automated tag suggestion
Katakis, I., Tsoumakas, G., and Vlahavas, I. (2008) · 2008
Cited alongside, same era.
A stick-breaking likelihood for categorical data analysis with latent Gaussian models
Khan, M. E., Mohamed, S., Marlin, B. M., and Murphy, K. P. (2012) · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
Mnih, A. and Teh, Y. W. (2012) · 2012
Cited alongside, same era.
Scalable Bayesian modelling of paired symbols
Paquet, U., Koenigstein, N., and Winther, O. (2012) · 2012
Cited alongside, same era.
Distributed training of large-scale logistic models
Gopal, S. and Yang, Y. (2013) · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Fast and robust neural network joint models for statistical machine translation
Devlin, J., Zbib, R., Huang, Z., Lamar, T., Schwartz, R., and Makhoul, J. (2014) · 2014
Later among the works it cites.
Glove: Global Vectors for Word Representation
Pennington, J., Socher, R., and Manning, C. (2014) · 2014
Later among the works it cites.
Deep networks with large output spaces
Vijayanarasimhan, S., Shlens, J., Monga, R., and Yagnik, J. (2014) · 2014
Later among the works it cites.
Sparse local embeddings for extreme multi-label classification
Bhatia, K., Jain, H., Kar, P., Varma, M., and Jain, P. (2015) · 2015
Later among the works it cites.
Blackout: Speeding up recurrent neural network language models with very large vocabularies
Ji, S., Vishwanathan, S. V. N., Satish, N., Anderson, M. J., and Dubey, P. (2015) · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013) · 2013
Cited alongside, same era.
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Closest in time.