Fetching the paper…
Reading the bibliography…
Categorical distributions are ubiquitous in machine learning, e.g., in classification, language models, and recommendation systems.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: A series of lectures
Gumbel, E. J · 1954
Earlier work this paper cites.
Modeling the choice of residential location
McFadden, D · 1978
Earlier work this paper cites.
On the multistage Bayes classifier
Kurzynski, M · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Bayesian analysis of binary and polychotomous response data
Albert, J. H. and Chib, S · 1993
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K · 1999
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Beal, M. J · 2003
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Bengio, Y. and Sénécal, J.-S · 2003
Earlier work this paper cites.
Bayesian data analysis
Gelman, A., Carlin, J. B., Stern, H. S., and Rubin, D. B · 2003
Earlier work this paper cites.
The multiple multiplicative factor model for collaborative filtering
Marlin, B. M. and Zemel, R. S · 2004
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, F. and Bengio, Y · 2005
Earlier work this paper cites.
Contrastive estimation: Training log-linear models on unlabeled data
Smith, N. A. and Jason, E · 2005
Earlier work this paper cites.
Neural probabilistic language models
Bengio, Y., Schwenk, H., Senécal, J.-S., Morin, F., and Gauvain, J.-L · 2006
Earlier work this paper cites.
Variational Bayesian multinomial probit regression with Gaussian process priors
Girolami, M. and Rogers, S · 2006
Earlier work this paper cites.
A correlated topic model of Science
Blei, D. M. and Lafferty, J. D · 2007
Earlier work this paper cites.
Efficient bounds for the softmax and applications to approximate inference in hybrid models
Bouchard, G · 2007
Earlier work this paper cites.
Multilabel text classification for automated tag suggestion
Katakis, I., Tsoumakas, G., and Vlahavas, I · 2008
Earlier work this paper cites.
Efficient pairwise multilabel classification for large-scale problems in the legal domain
Mencia, E. L. and Furnkranz, J · 2008
Cited alongside, same era.
Effective and efficient multilabel classification in domains with large number of labels
Tsoumakas, G., Katakis, I., and Vlahavas, I · 2008
Cited alongside, same era.
Conditional probability tree estimation analysis and algorithms
Beygelzimer, A., Langford, J., Lifshits, Y., Sorkin, G. B., and Strehl, L · 2009
Cited alongside, same era.
Bayes optimal multilabel classification via probabilistic classifier chains
Dembczyński, K., Cheng, W., and Hüllermeier, E · 2010
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Doubly stochastic variational Bayes for non-conjugate inference
Titsias, M. K. and Lázaro-Gredilla, M · 2014
Later among the works it cites.
Deep networks with large output spaces
Vijayanarasimhan, S., Shlens, J., Monga, R., and Yagnik, J · 2014
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Later among the works it cites.
Sparse local embeddings for extreme multi-label classification
Bhatia, K., Jain, H., Kar, P., Varma, M., and Jain, P · 2015
Later among the works it cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Later among the works it cites.
Efficient exact gradient update for training deep networks with very large sparse targets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
A stick-breaking likelihood for categorical data analysis with latent Gaussian models
Khan, M. E., Mohamed, S., Marlin, B. M., and Murphy, K. P · 2012
Cited alongside, same era.
Lecture 6.5-RMSPROP: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Distributed training of large-scale logistic models
Gopal, S. and Yang, Y · 2013
Cited alongside, same era.
Stochastic variational inference
Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J · 2013
Cited alongside, same era.
Hidden factors and hidden topics: understanding rating dimensions with review text
McAuley, J. and Leskovec, J · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Cited alongside, same era.
Vincent, P., de Brébisson, A., and Bouthillier, X · 2015
Later among the works it cites.
Importance weighted autoencoders
Burda, Y., Grosse, R., and Salakhutdinov, R · 2016
Later among the works it cites.
An exploration of softmax alternatives belonging to the spherical loss family
de Brébisson, A. and Vincent, P · 2016
Later among the works it cites.
Blackout: Speeding up recurrent neural network language models with very large vocabularies
Ji, S., Vishwanathan, S. V. N., Satish, N., Anderson, M. J., and Dubey, P · 2016
Later among the works it cites.
One-vs-each approximation to softmax for scalable estimation of probabilities
Titsias, M. K · 2016
Later among the works it cites.
Variational inference: A review for statisticians
Blei, D., Kucukelbir, A., and McAuliffe, J · 2017
Later among the works it cites.
Complementary sum sampling for likelihood approximation in large scale classification
Botev, A., Zheng, B., and Barber, D · 2017
Later among the works it cites.
Efficient softmax approximation for GPUs
Grave, E., Joulin, A., Cissé, M., Grangier, D., and Jégrou, H · 2017
Later among the works it cites.
Automatic differentiation variational inference
Kucukelbir, A., Tran, D., Ranganath, R., Gelman, A., and Blei, D. M · 2017
Later among the works it cites.
Fast amortized inference and learning in log-linear models with randomly perturbed nearest neighbor search
Mussmann, S., Levy, D., and Ermon, S · 2017
Later among the works it cites.
DS-MLR: exploiting double separability for scaling up distributed multinomial logistic regression
Raman, P., Matsushima, S., Zhang, X., Yun, H., and Vishwanathan, S. V. N · 2017
Later among the works it cites.
Sticking the landing: Simple, lower-variance gradient estimators for variational inference
Roeder, G., Wu, Y., and Duvenaud, D · 2017
Later among the works it cites.
Unbiased scalable softmax optimization
Fagan, F. and Iyengar, G · 2018
Closest in time.