Fetching the paper…
Reading the bibliography…
What is a systematic way to efficiently apply a wide spectrum of advanced ML programs to industrial scale problems, using Big Models (up to 100s of billions of parameters) on Big Data (up to terabytes or petabytes)? Modern parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-synchronous processing paradigm popularized by MapReduce, or even specialized graph-based execution that relies on graph representations of ML programs.
Distance metric learning with application to clustering with side-information
E. P. Xing, M. I. Jordan, S. Russell, and A. Y. Ng · 2002
Earlier work this paper cites.
Finding scientific topics
T. L. Griffiths and M. Steyvers · 2004
Earlier work this paper cites.
Information-theoretic metric learning
J. V. Davis, B. Kulis, P. Jain, S. Sra, and I. S. Dhillon · 2007
Earlier work this paper cites.
Large-scale parallel collaborative filtering for the netflix prize
Y. Zhou, D. Wilkinson, R. Schreiber, and R. Pan · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
M. Zinkevich, J. Langford, and A. J. Smola · 2009
Earlier work this paper cites.
Pregel: a system for large-scale graph processing
G. Malewicz, M. H. Austern, A. J. Bik, J. C. Dehnert, I. Horn, N. Leiser, and G. Czajkowski · 2010
Earlier work this paper cites.
Piccolo: building fast, distributed programs with partitioned tables
R. Power and J. Li · 2010
Earlier work this paper cites.
Spark: cluster computing with working sets
M. Zaharia, M. Chowdhury, M. J. Franklin, S. Shenker, and I. Stoica · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
Parallel coordinate descent for l1-regularized loss minimization
J. K. Bradley, A. Kyrola, D. Bickson, and C. Guestrin · 2011
Earlier work this paper cites.
Smoothing proximal gradient method for general structured sparse learning
X. Chen, Q. Lin, S. Kim, J. Carbonell, and E. Xing · 2011
Cited alongside, same era.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Ré, and S. J. Wright · 2011
Cited alongside, same era.
Priter: A distributed framework for prioritized iterative computations
Y. Zhang, Q. Gao, L. Gao, and C. Wang · 2011
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. Le, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Ng · 2012
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Q. Le, M. Ranzato, R. Monga, M. Devin, K. Chen, G. Corrado, J. Dean, and A. Ng · 2012
Cited alongside, same era.
Distributed GraphLab: A Framework for Machine Learning and Data Mining in the Cloud
More effective distributed ml via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, J.-K. Kim, S. Lee, P. B. Gibbons, G. Gibson, G. R. Ganger, and E. P. Xing · 2013
Closest in time.
Stochastic variational inference
M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley · 2013
Closest in time.
Parallel markov chain monte carlo for nonparametric mixture models
S. A. Williamson, A. Dubey, and E. P. Xing · 2013
Closest in time.
Priter: A distributed framework for prioritizing iterative computations
Y. Zhang, Q. Gao, L. Gao, and C. Wang · 2013
Closest in time.
Fugue: Slow-worker-agnostic distributed learning for big models on big data
A. Kumar, A. Beutel, Q. Ho, and E. P. Xing · 2014
Closest in time.
On model parallelism and scheduling strategies for distributed machine learning
S. Lee, J. K. Kim, X. Zheng, Q. Ho, G. Gibson, and E. P. Xing · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Low, J. Gonzalez, A. Kyrola, D. Bickson, C. Guestrin, and J. M. Hellerstein · 2012
Cited alongside, same era.
Parallel coordinate descent methods for big data optimization
P. Richtárik and M. Takáč · 2012
Cited alongside, same era.
Feature clustering for accelerating parallel coordinate descent
C. Scherrer, A. Tewari, M. Halappanavar, and D. Haglin · 2012
Cited alongside, same era.
Hadoop: The definitive guide
T. White · 2012
Cited alongside, same era.
Scalable coordinate descent approaches to parallel matrix factorization for recommender systems
H.-F. Yu, C.-J. Hsieh, S. Si, and I. Dhillon · 2012
Cited alongside, same era.
Ad click prediction: a view from the trenches
H. B. M. et. al · 2013
Cited alongside, same era.
Fugue: Slow-worker-agnostic distributed learning for big models on big data
A. Kumar, A. Beutel, Q. Ho, and E. P. Xing
Cited in the paper.
Closest in time.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Closest in time.
Towards topic modeling for big data
Y. Wang, X. Zhao, Z. Sun, H. Yan, L. Wang, Z. Jin, L. Wang, Y. Gao, J. Zeng, Q. Yang, et al · 2014
Closest in time.
High-performance distributed ml at scale through parameter server consistency models
W. Dai, A. Kumar, J. Wei, Q. Ho, G. Gibson, and E. P. Xing · 2015
Closest in time.
Lightlda: Big topic models on modest compute clusters
J. Yuan, F. Gao, Q. Ho, W. Dai, J. Wei, X. Zheng, E. P. Xing, T.-Y. Liu, and W.-Y. Ma · 2015
Closest in time.