Fetching the paper…
Reading the bibliography…
The Softmax function is ubiquitous in machine learning, multiple previous works suggested faster alternatives for it.
Note on a method for calculating corrected sums of squares and products
B. P. Welford · 1962
Earlier work this paper cites.
Classes for fast maximum entropy training
Joshua Goodman · 2001
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Yoshua Bengio and Jean-Sébastien Sénécal · 2003
Earlier work this paper cites.
Extensions of recurrent neural network language model
T. Mikolov, S. Kombrink, L. Burget, J. Černocký, and S. Khudanpur · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
A Fast and Simple Algorithm for Training Neural Probabilistic Language Models
A. Mnih and Y. Whye Teh · 2012
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Fast and robust neural network joint models for statistical machine translation
Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, and John Makhoul · 2014
Cited alongside, same era.
Strategies for Training Large Vocabulary Neural Language Models
W. Chen, D. Grangier, and M. Auli · 2015
Cited alongside, same era.
BlackOut: Speeding up Recurrent Neural Network Language Models With Very Large Vocabularies
S. Ji, S. V. N. Vishwanathan, N. Satish, M. J. Anderson, and P. Dubey · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Chainer: a next-generation open source framework for deep learning
Seiya Tokui, Kenta Oono, Shohei Hido, and Justin Clayton · 2015
Later among the works it cites.
Efficient softmax approximation for GPUs
E. Grave, A. Joulin, M. Cissé, D. Grangier, and H. Jégou · 2016
Later among the works it cites.
Cntk: Microsoft’s open-source deep-learning toolkit
Frank Seide and Amit Agarwal · 2016
Later among the works it cites.
Svd-softmax: Fast softmax approximation on large vocabulary neural networks
Kyuhong Shim, Minjae Lee, Iksoo Choi, Yoonho Boo, and Wonyong Sung · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Cited alongside, same era.
Polytomous logistic regression
Engel J
Cited in the paper.