Fetching the paper…
Reading the bibliography…
Scaling multinomial logistic regression to datasets with very large number of data points and classes is challenging.
A stochastic approximation method
H. E. Robbins and S. Monro · 1951
Earlier work this paper cites.
Nonlinear programming
D. P. Bertsekas · 1999
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
A. Nedić and D. Bertsekas · 2001
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
H. Kushner and G. G. Yin · 2003
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Earlier work this paper cites.
Efficient bounds for the softmax function, applications to inference in hybrid models
G. Bouchard · 2007
Earlier work this paper cites.
Slow learners are fast
M. Zinkevich, J. Langford, and A. J. Smola · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
The youtube video recommendation system
J. Davidson, B. Liebald, J. Liu, P. Nandy, T. Van Vleet, U. Gargi, S. Gupta, Y. He, M. Lambert, B. Livingston, et al · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Cited alongside, same era.
The tradeoffs of large-scale learning
L. Bottou and O. Bousquet · 2011
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Cited alongside, same era.
Large-scale matrix factorization with distributed stochastic gradient descent
R. Gemulla, E. Nijkamp, P. J. Haas, and Y. Sismanis · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Cited alongside, same era.
A fast parallel stochastic gradient method for matrix factorization in shared memory systems
Y. Zhuang, Y.-C. Juan, and C.-J. Lin · 2013
Later among the works it cites.
Doubly Separable Models
H. Yun · 2014
Later among the works it cites.
Ranking via robust binary classification
H. Yun, P. Raman, and S. Vishwanathan · 2014
Later among the works it cites.
An asynchronous parallel stochastic coordinate descent algorithm
J. Liu, S. J. Wright, C. Ré, V. Bittorf, and S. Sridhar · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Later among the works it cites.
Distributed machine learning via sufficient factor broadcasting
P. Xie, J. K. Kim, Y. Zhou, Q. Ho, A. Kumar, Y. Yu, and E. P. Xing · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. D. Zeiler · 2012
Cited alongside, same era.
Distributed training of large-scale logistic models
S. Gopal and Y. Yang · 2013
Cited alongside, same era.
Parameter server for distributed machine learning
M. Li, L. Zhou, Z. Yang, A. Li, F. Xia, D. G. Andersen, and A. Smola · 2013
Cited alongside, same era.
Nomad: Non-locking, stochastic multi-machine algorithm for asynchronous and decentralized matrix completion
H. Yun, H.-F. Yu, C.-J. Hsieh, S. Vishwanathan, and I. Dhillon · 2013
Cited alongside, same era.
Petuum: a new platform for distributed machine learning on big data
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu · 2015
Later among the works it cites.
Asaga: asynchronous parallel saga
R. Leblond, F. Pedregosa, and S. Lacoste-Julien · 2016
Closest in time.
Pd-sparse : A primal and dual sparse approach to extreme multiclass and multilabel classification
I. E.-H. Yen, X. Huang, P. Ravikumar, K. Zhong, and I. Dhillon · 2016
Closest in time.