Fetching the paper…
Reading the bibliography…
Recent neural network and language models rely on softmax distributions with an extremely large number of categories.
The maximum concurrent flow problem
Shahrokhi, Farhad and Matula, David W · 1990
Earlier work this paper cites.
Customer satisfaction, customer retention, and market share
Rust, Roland T and Zahorik, Anthony J · 1993
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Bengio, Yoshua, Senécal, Jean-Sébastien, et al · 2003
Earlier work this paper cites.
Convex optimization
Boyd, Stephen and Vandenberghe, Lieven · 2004
Earlier work this paper cites.
Sparse multinomial logistic regression: Fast algorithms and generalization bounds
Krishnapuram, Balaji, Carin, Lawrence, Figueiredo, Mario AT, and Hartemink, Alexander J · 2005
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Bengio, Yoshua and Senécal, Jean-Sébastien · 2008
Earlier work this paper cites.
Essential medical statistics
Kirkwood, Betty R and Sterne, Jonathan AC · 2010
Earlier work this paper cites.
Incremental proximal methods for large scale convex optimization
Bertsekas, Dimitri P · 2011
Earlier work this paper cites.
Lacoste-Julien, Simon, Schmidt, Mark, and Bach, Francis · 2012
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Mnih, Andriy and Teh, Yee Whye · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, Ciprian, Mikolov, Tomas, Schuster, Mike, Ge, Qi, Brants, Thorsten, Koehn, Phillipp, and Robinson, Tony · 2013
Cited alongside, same era.
Distributed training of large-scale logistic models
Gopal, Siddharth and Yang, Yiming · 2013
Cited alongside, same era.
Stochastic proximal iteration: a non-asymptotic improvement upon stochastic gradient descent
Ryu, Ernest K and Boyd, Stephen · 2014
Cited alongside, same era.
When and why are log-linear models self-normalizing?
Andreas, Jacob and Klein, Dan · 2015
Cited alongside, same era.
An exploration of softmax alternatives belonging to the spherical loss family
de Brébisson, Alexandre and Vincent, Pascal · 2015
Cited alongside, same era.
Efficient softmax approximation for GPUs
Grave, Edouard, Joulin, Armand, Cissé, Moustapha, Grangier, David, and Jégou, Hervé · 2016
Later among the works it cites.
Jernite, Yacine, Choromanska, Anna, Sontag, David, and LeCun, Yann · 2016
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, Chris J, Mnih, Andriy, and Teh, Yee Whye · 2016
Later among the works it cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Martins, André FT and Astudillo, Ramón Fernandez · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ji, Shihao, Vishwanathan, SVN, Satish, Nadathur, Anderson, Michael J, and Dubey, Pradeep · 2015
Cited alongside, same era.
LSHTC: A benchmark for large-scale text classification
Partalas, Ioannis, Kosmopoulos, Aris, Baskiotis, Nicolas, Artieres, Thierry, Paliouras, George, Gaussier, Eric, Androutsopoulos, Ion, Amini, Massih-Reza, and Galinari, Patrick · 2015
Cited alongside, same era.
Implicit stochastic approximation
Toulis, Panos and Airoldi, Edoardo M · 2015
Cited alongside, same era.
Efficient exact gradient update for training deep networks with very large sparse targets
Vincent, Pascal, de Brébisson, Alexandre, and Bouthillier, Xavier · 2015
Cited alongside, same era.
Logarithmic time one-against-some
Daume III, Hal, Karampatziakis, Nikos, Langford, John, and Mineiro, Paul · 2016
Cited alongside, same era.
Raman, Parameswaran, Matsushima, Shin, Zhang, Xinhua, Yun, Hyokun, and Vishwanathan, SVN · 2016
Later among the works it cites.
One-vs-each approximation to softmax for scalable estimation of probabilities
Titsias, Michalis K · 2016
Later among the works it cites.
Towards stability and optimality in stochastic gradient descent
Toulis, Panos, Tran, Dustin, and Airoldi, Edo · 2016
Later among the works it cites.
Accelerating stochastic composition optimization
Wang, Mengdi, Liu, Ji, and Fang, Ethan · 2016
Later among the works it cites.
Aggressive sampling for multi-class to binary reduction with applications to text classification
Joshi, Bikash, Amini, Massih-Reza, Partalas, Ioannis, Iutzeler, Franck, and Maximov, Yury · 2017
Later among the works it cites.
Augment and reduce: Stochastic inference for large categorical distributions
Ruiz, Francisco JR, Titsias, Michalis K, Dieng, Adji B, and Blei, David M · 2018
Closest in time.