Fetching the paper…
Reading the bibliography…
Motivated by recent developments in serverless systems for large-scale computation as well as improvements in scalable randomized matrix algorithms, we develop OverSketched Newton, a randomized Hessian-based optimization algorithm to solve large-scale convex optimization problems in serverless systems.
J. Levin, “Note on convergence of minres,” Multivariate behavioral research
1988
Earlier work this paper cites.
J. R. Shewchuk et al
1994
Earlier work this paper cites.
R. A. van de Geijn and J. Watts, “Summa: Scalable universal matrix multiplication algorithm,” tech. rep., 1995
1995
Earlier work this paper cites.
New York, NY, USA: Cambridge University Press, 2004
S. Boyd and L. Vandenberghe, Convex Optimization · 2004
Earlier work this paper cites.
Springer Science & Business Media, 2006
J. Nocedal and S. Wright, Numerical optimization · 2006
Earlier work this paper cites.
J. Dean and S. Ghemawat, “Mapreduce: Simplified data processing on large clusters,” Commun. ACM
2008
Earlier work this paper cites.
T. Hoefler, T. Schneider, and A. Lumsdaine, “Characterizing the influence of system noise on large-scale applications by simulation,” in Proc. of the ACM/IEEE Int. Conf. for High Perf. Comp., Networking, Storage and Analysis
2010
Earlier work this paper cites.
M. Zaharia, M. Chowdhury, M. J. Franklin, S. Shenker, and I. Stoica, “Spark: Cluster computing with working sets,” in Proceedings of the 2Nd USENIX Conference on Hot Topics in Cloud Computing
2010
Earlier work this paper cites.
Foundations and Trends in Machine Learning, Boston: NOW Publishers, 2011
M. W. Mahoney, Randomized algorithms for matrices and data · 2011
Earlier work this paper cites.
C.-C. Chang and C.-J. Lin, “Libsvm: a library for support vector machines,” ACM transactions on intelligent systems and technology (TIST)
2011
Earlier work this paper cites.
E. Solomonik and J. Demmel, “Communication-optimal parallel 2.5D matrix multiplication and LU factorization algorithms,” in Proceedings of the 17th International Conference on Parallel Processing
2011
Earlier work this paper cites.
American Mathematical Soc., 2012
T. Tao, Topics in random matrix theory · 2012
Earlier work this paper cites.
J. Dean and L. A. Barroso, “The tail at scale,” Commun. ACM
2013
Earlier work this paper cites.
J. Nelson and H. L. Nguyen, “Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings,” in 2013 IEEE 54th Annual Symposium on Foundations of Computer Science
2013
Earlier work this paper cites.
D. P. Woodruff, “Sketching as a tool for numerical linear algebra,” Found. Trends Theor. Comput. Sci
2014
Earlier work this paper cites.
O. Shamir, N. Srebro, and T. Zhang, “Communication-efficient distributed optimization using an approximate Newton-type method,” in Proceedings of the 31st International Conference on International Conference on Machine Learning - Volume 32
2014
Earlier work this paper cites.
Springer Publishing Company, Incorporated, 1 ed., 2014
Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course · 2014
Earlier work this paper cites.
Y. Zhang and X. Lin, “Disco: Distributed optimization for self-concordant empirical loss,” in Proceedings of the 32nd International Conference on Machine Learning
2015
Earlier work this paper cites.
S. J. Reddi, A. Hefny, S. Sra, B. Pöczos, and A. Smola, “On variance reduction in stochastic gradient descent and its asynchronous variants,” in Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
A. Gittens, A. Devarakonda, E. Racah, M. Ringenburg, L. Gerhardt, J. Kottalam, J. Liu, K. Maschhoff, S. Canon, J. Chhugani, et al
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” Siam Review
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Wang, F. Roosta-Khorasani, P. Xu, and M. W. Mahoney, “GIANT: Globally improved approximate Newton method for distributed optimization,” in Advances in Neural Information Processing Systems
2018
Later among the works it cites.
C.-H. Fang, S. B. Kylasa, F. Roosta-Khorasani, M. W. Mahoney, and A. Grama, “Distributed Second-order Convex Optimization,” ArXiv e-prints
2018
Later among the works it cites.
V. Gupta, S. Wang, T. Courtade, and K. Ramchandran, “Oversketch: Approximate matrix multiplication for the cloud,” IEEE International Conference on Big Data, Seattle, WA, USA
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
J. Spillner, C. Mateos, and D. A. Monge, “Faaster, better, cheaper: The prospect of serverless scientific computing and HPC,” in Latin American High Performance Computing Conference
2017
Cited alongside, same era.
E. Jonas, Q. Pu, S. Venkataraman, I. Stoica, and B. Recht, “Occupy the cloud: distributed computing for the 99%,” in Proceedings of the 2017 Symposium on Cloud Computing
2017
Cited alongside, same era.
2017
Cited alongside, same era.
P. Xu, F. Roosta, and M. W. Mahoney, “Newton-type methods for non-convex optimization under inexact hessian information,” 2017
2017
Cited alongside, same era.
M. Pilanci and M. J. Wainwright, “Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence,” SIAM Jour. on Opt
2017
Cited alongside, same era.
R. Tandon, Q. Lei, A. G. Dimakis, and N. Karampatziakis, “Gradient coding: Avoiding stragglers in distributed learning,” in Proceedings of the 34th International Conference on Machine Learning
2017
Cited alongside, same era.
Q. Yu, M. Maddah-Ali, and S. Avestimehr, “Polynomial codes: an optimal design for high-dimensional coded matrix multiplication,” in Advances in Neural Inf. Processing Systems 30
2017
Cited alongside, same era.
2018
Later among the works it cites.
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran, “Speeding up distributed machine learning using codes,” IEEE Transactions on Information Theory
2018
Later among the works it cites.
T. Baharav, K. Lee, O. Ocal, and K. Ramchandran, “Straggler-proofing massive-scale distributed matrix multiplication with d-dimensional product codes,” in IEEE Int. Sym. on Information Theory (ISIT)
2018
Later among the works it cites.
C. Duenner, A. Lucchi, M. Gargiani, A. Bian, T. Hofmann, and M. Jaggi, “A distributed second-order algorithm you can trust,” in Proceedings of the 35th International Conference on Machine Learning
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.
H. Wang, D. Niu, and B. Li, “Distributed machine learning with a serverless architecture,” in IEEE INFOCOM 2019-IEEE Conference on Computer Communications
2019
Closest in time.
J. Carreira, P. Fonseca, A. Tumanov, A. Zhang, and R. Katz, “Cirrus: a serverless framework for end-to-end ml workflows,” in Proceedings of the ACM Symposium on Cloud Computing
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
V. Gupta, D. Carrano, Y. Yang, V. Shankar, T. Courtade, and K. Ramchandran, “Serverless straggler mitigation using local error-correcting codes,” IEEE International Conference on Distributed Computing and Systems (ICDCS), Singapore
2020
Closest in time.
Technavio, “Serverless architecture market by end-users and geography - global forecast 2019-2023.” https://www.technavio.com/report/serverless-architecture-market-industry-analysis
2023
Closest in time.