Fetching the paper…
Reading the bibliography…
We study the problem of distributed multi-task learning with shared representation, where each machine aims to learn a separate, but related, task in an unknown shared low-dimensional subspaces, i.e.
An algorithm for quadratic programming
M. Frank and P. Wolfe · 1956
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate ≀ ( 1 / k 2 ) {\cal o}(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Database of homology-derived protein structures and the structural meaning of sequence alignment
C. Sander and R. Schneider · 1991
Earlier work this paper cites.
Hierarchical bayes conjoint analysis: Recovery of partworth heterogeneity from reduced experimental designs
P. J. Lenk, W. S. DeSarbo, P. E. Green, and M. R. Young · 1996
Earlier work this paper cites.
Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Greed is good: Algorithmic results for sparse approximation
J. A. Tropp · 2004
Earlier work this paper cites.
A framework for learning predictive structures from multiple tasks and unlabeled data
R. K. Ando and T. Zhang · 2005
Earlier work this paper cites.
Uncovering shared structures in multiclass classification
Y. Amit, M. Fink, N. Srebro, and S. Ullman · 2007
Earlier work this paper cites.
Multi-task learning for classification with dirichlet process priors
Y. Xue, X. Liao, L. Carin, and B. Krishnapuram · 2007
Earlier work this paper cites.
Dimension reduction and coefficient estimation in multivariate linear regression
M. Yuan, A. Ekici, Z. Lu, and R. Monteiro · 2007
Earlier work this paper cites.
Convex multi-task feature learning
A. Argyriou, T. Evgeniou, and M. Pontil · 2008
Earlier work this paper cites.
The tradeoffs of large scale learning
O. Bousquet and L. Bottou · 2008
Earlier work this paper cites.
Semantic annotation and retrieval of music and sound effects
D. Turnbull, L. Barrington, D. Torres, and G. Lanckriet · 2008
Earlier work this paper cites.
An accelerated gradient method for trace norm minimization
S. Ji and J. Ye · 2009
Earlier work this paper cites.
Feature hashing for large scale multitask learning
K. Weinberger, A. Dasgupta, J. Langford, A. Smola, and J. Attenberg · 2009
Earlier work this paper cites.
A singular value thresholding algorithm for matrix completion
J.-F. Cai, E. J. Candès, and Z. Shen · 2010
Cited alongside, same era.
Multi-task learning for boosting with application to web search ranking
O. Chapelle, P. Shivaswamy, S. Vadrevu, K. Weinberger, Y. Zhang, and B. Tseng · 2010
Cited alongside, same era.
Tree-guided group lasso for multi-task regression with structured sparsity
S. Kim and E. P. Xing · 2010
Cited alongside, same era.
Distributed stochastic subgradient projection algorithms for convex optimization
S. S. Ram, A. Nedić, and V. V. Veeravalli · 2010
Cited alongside, same era.
Trading accuracy for sparsity in optimization problems with sparsity constraints
S. Shalev-Shwartz, N. Srebro, and T. Zhang · 2010
Cited alongside, same era.
Optimization with sparsity-inducing penalties
F. Bach, R. Jenatton, J. Mairal, and G. Obozinski · 2011
Revisiting frank-wolfe: Projection-free sparse convex optimization
M. Jaggi · 2013
Later among the works it cites.
Low-rank matrix completion using alternating minimization
P. Jain, P. Netrapalli, and S. Sanghavi · 2013
Later among the works it cites.
Excess risk bounds for multitask learning with trace norm regularization
A. Maurer and M. Pontil · 2013
Later among the works it cites.
Multi-task learning in deep neural networks for improved phoneme recognition
M. L. Seltzer and J. Droppo · 2013
Later among the works it cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Y. Zhang, J. C. Duchi, M. I. Jordan, and M. J. Wainwright · 2013
Later among the works it cites.
Modeling disease progression via multi-task learning
J. Zhou, J. Liu, V. A. Narayan, and J. Ye · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scaling up machine learning: Parallel and distributed approaches
R. Bekkerman, M. Bilenko, and J. Langford · 2011
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. P. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Cited alongside, same era.
Natural language processing (almost) from scratch
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa · 2011
Cited alongside, same era.
Large-scale convex minimization with a low-rank constraint
S. Shalev-Shwartz, A. Gonen, and O. Shamir · 2011
Cited alongside, same era.
Fast global convergence of gradient methods for high-dimensional statistical recovery
A. Agarwal, S. Negahban, and M. J. Wainwright · 2012
Cited alongside, same era.
Distributed learning, communication complexity and privacy
M.-F. Balcan, A. Blum, S. Fine, and Y. Mansour · 2012
Cited alongside, same era.
Later among the works it cites.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan · 2014
Later among the works it cites.
Scalable multitask representation learning for scene classification
M. Lapin, B. Schiele, and M. Hein · 2014
Later among the works it cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Later among the works it cites.
Distributed stochastic optimization and learning
O. Shamir and N. Srebro · 2014
Later among the works it cites.
Communication efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Later among the works it cites.
A distributed frank-wolfe algorithm for communication-efficient sparse learning
A. Bellet, Y. Liang, A. B. Garakani, M.-F. Balcan, and F. Sha · 2015
Later among the works it cites.
On the global linear convergence of frank-wolfe optimization variants
S. Lacoste-Julien and M. Jaggi · 2015
Later among the works it cites.
Communication-efficient sparse regression: a one-shot approach
J. D. Lee, Y. Sun, Q. Liu, and J. E. Taylor · 2015
Later among the works it cites.
Communication-efficient distributed optimization of self-concordant empirical loss
Y. Zhang and L. Xiao · 2015
Later among the works it cites.