Fetching the paper…
Reading the bibliography…
Matrix-parametrized models, including multiclass logistic regression and sparse coding, are used in machine learning (ML) applications ranging from computer vision to computational biology.
Information processing in dynamical systems: Foundations of harmony theory
P. Smolensky · 1986
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. P. Bertsekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Nonlinear programming
D. P. Bertsekas · 1999
Earlier work this paper cites.
Learning the parts of objects by non-negative matrix factorization
D. D. Lee and H. S. Seung · 1999
Earlier work this paper cites.
Distance metric learning with application to clustering with side-information
E. P. Xing, M. I. Jordan, S. Russell, and A. Y. Ng · 2002
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
K. Q. Weinberger, J. Blitzer, and L. K. Saul · 2005
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
M. Yuan and Y. Lin · 2006
Earlier work this paper cites.
Distributed decision-tree induction in peer-to-peer systems
K. Bhaduri, R. Wolff, C. Giannella, and H. Kargupta · 2008
Earlier work this paper cites.
Mapreduce: simplified data processing on large clusters
J. Dean and S. Ghemawat · 2008
Earlier work this paper cites.
A dual coordinate descent method for large-scale linear svm
C.-J. Hsieh, K.-W. Chang, C.-J. Lin, S. S. Keerthi, and S. Sundararajan · 2008
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
A. Beck and M. Teboulle · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
A local asynchronous distributed privacy preserving feature selection algorithm for large peer-to-peer networks
K. Das, K. Bhaduri, and H. Kargupta · 2010
Cited alongside, same era.
Pregel: a system for large-scale graph processing
G. Malewicz, M. H. Austern, A. J. Bik, J. C. Dehnert, I. Horn, N. Leiser, and G. Czajkowski · 2010
Cited alongside, same era.
Locality-constrained linear coding for image classification
J. Wang, J. Yang, K. Yu, F. Lv, T. Huang, and Y. Gong · 2010
Cited alongside, same era.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Cited alongside, same era.
Scalable inference in latent variable models
A. Ahmed, M. Aly, J. Gonzalez, S. Narayanamurthy, and A. J. Smola · 2012
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Stochastic dual coordinate ascent methods for regularized loss
S. Shalev-Shwartz and T. Zhang · 2013
Later among the works it cites.
Replicated data consistency explained through baseball
D. Terry · 2013
Later among the works it cites.
Trading computation for communication: Distributed stochastic dual coordinate ascent
T. Yang · 2013
Later among the works it cites.
Project adam: building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Closest in time.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan · 2014
Closest in time.
On model parallelization and scheduling strategies for distributed machine learning
S. Lee, J. K. Kim, X. Zheng, Q. Ho, G. A. Gibson, and E. P. Xing · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Powergraph: distributed graph-parallel computation on natural graphs
J. E. Gonzalez, Y. Low, H. Gu, D. Bickson, and C. Guestrin · 2012
Cited alongside, same era.
Large-scale distributed non-negative sparse coding and sparse dictionary learning
V. Sindhwani and A. Ghoting · 2012
Cited alongside, same era.
Communication/computation tradeoffs in consensus-based distributed optimization
K. Tsianos, S. Lawlor, and M. G. Rabbat · 2012
Cited alongside, same era.
Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing
M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M. J. Franklin, S. Shenker, and I. Stoica · 2012
Cited alongside, same era.
Distributed training of large-scale logistic models
S. Gopal and Y. Yang · 2013
Cited alongside, same era.
More effective distributed ml via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. Xing · 2013
Cited alongside, same era.
Closest in time.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Closest in time.
Communication efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Closest in time.
High-performance distributed ml at scale through parameter server consistency models
W. Dai, A. Kumar, J. Wei, Q. Ho, G. Gibson, and E. P. Xing · 2015
Closest in time.
Passcode: Parallel asynchronous stochastic dual co-ordinate descent
C.-J. Hsieh, H.-F. Yu, and I. S. Dhillon · 2015
Closest in time.
Malt: distributed data-parallelism for existing ml applications
H. Li, A. Kadav, E. Kruus, and C. Ungureanu · 2015
Closest in time.
Lshtc: A benchmark for large-scale text classification
I. Partalas, A. Kosmopoulos, N. Baskiotis, T. Artieres, G. Paliouras, E. Gaussier, I. Androutsopoulos, M.-R. Amini, and P. Galinari · 2015
Closest in time.