Fetching the paper…
Reading the bibliography…
We study robust distributed learning that involves minimizing a non-convex loss function with saddle points.
The Byzantine generals problem
Leslie Lamport, Robert Shostak, and Marshall Pease · 1982
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Nemirovskii, David Borisovich Yudin, and Edgar Ronald Dawson · 1983
Earlier work this paper cites.
Random generation of combinatorial structures from a uniform distribution
Mark R Jerrum, Leslie G Valiant, and Vijay V Vazirani · 1986
Earlier work this paper cites.
Distributed Algorithms
Nancy A. Lynch · 1996
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Dimitri P. Bertsekas and John N. Tsitsiklis · 2000
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Gradient diversity: a key ingredient for scalable distributed learning
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2007
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Robust statistics
Peter J Huber · 2011
Earlier work this paper cites.
Learning sparsely used overcomplete dictionaries
Alekh Agarwal, Animashree Anandkumar, Prateek Jain, Praneeth Netrapalli, and Rashish Tandon · 2014
Earlier work this paper cites.
Statistical guarantees for the EM algorithm: From population to sample-based analysis
Sivaraman Balakrishnan, Martin J. Wainwright, and Bin Yu · 2014
Earlier work this paper cites.
First-order methods of smooth convex optimization with inexact oracle
Olivier Devolder, François Glineur, and Yurii Nesterov · 2014
Earlier work this paper cites.
Jiashi Feng, Huan Xu, and Shie Mannor · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Earlier work this paper cites.
Robust regression via hard thresholding
Kush Bhatia, Prateek Jain, and Purushottam Kar · 2015
Earlier work this paper cites.
Phase retrieval via Wirtinger flow: Theory and algorithms
Emmanuel J Candes, Xiaodong Li, and Mahdi Soltanolkotabi · 2015
Earlier work this paper cites.
Yudong Chen and Martin J Wainwright · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Federated optimization: Distributed optimization beyond the datacenter
Jakub Konečnỳ, Brendan McMahan, and Daniel Ramage · 2015
Earlier work this paper cites.
Jason D Lee, Qihang Lin, Tengyu Ma, and Tianbao Yang · 2015
Earlier work this paper cites.
Geometric median and robust estimation in banach spaces
Stanislav Minsker et al · 2015
Earlier work this paper cites.
Complete dictionary recovery using nonconvex optimization
J. Sun, Q. Qu, and J. Wright · 2015
Cited alongside, same era.
Low-rank solutions of linear matrix equations via procrustes flow
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Benjamin Recht · 2015
Cited alongside, same era.
A nonconvex optimization framework for low rank matrix estimation
Tuo Zhao, Zhaoran Wang, and Han Liu · 2015
Cited alongside, same era.
Finding approximate local minima for nonconvex optimization in linear time
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2016
Cited alongside, same era.
Accelerated methods for non-convex optimization
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2016
Cited alongside, same era.
Distributed statistical machine learning in adversarial settings: Byzantine gradient descent
Yudong Chen, Lili Su, and Jiaming Xu · 2017
Later among the works it cites.
A trust region algorithm with a worst-case iteration complexity of ϵ − 3 / 2 \epsilon^{-3/2} for nonconvex optimization
Frank E Curtis, Daniel P Robinson, and Mohammadreza Samadi · 2017
Later among the works it cites.
Being robust (in high dimensions) can be practical
Ilias Diakonikolas, Gautam Kamath, Daniel M Kane, Jerry Li, Ankur Moitra, and Alistair Stewart · 2017
Later among the works it cites.
Gradient descent can take exponential time to escape saddle points
Simon S Du, Chi Jin, Jason D Lee, Michael I Jordan, Aarti Singh, and Barnabas Poczos · 2017
Later among the works it cites.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust estimators in high dimensions without the computational intractability
Ilias Diakonikolas, Gautam Kamath, Daniel M Kane, Jerry Li, Ankur Moitra, and Alistair Stewart · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Federated optimization: distributed machine learning for on-device intelligence
Jakub Konečnỳ, H Brendan McMahan, Daniel Ramage, and Peter Richtárik · 2016
Cited alongside, same era.
Agnostic estimation of mean and covariance
Kevin A Lai, Anup B Rao, and Santosh Vempala · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
The power of normalization: Faster evasion of saddle points
Kfir Y Levy · 2016
Cited alongside, same era.
Rong Ge, Chi Jin, and Yi Zheng · 2017
Later among the works it cites.
First-order methods almost always avoid saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2017
Later among the works it cites.
Robust sparse estimation tasks in high dimensions
Jerry Li · 2017
Later among the works it cites.
Federated learning: Collaborative machine learning without centralized training data
Brendan McMahan and Daniel Ramage · 2017
Later among the works it cites.
Distributed statistical estimation and rates of convergence in normal approximation
Stanislav Minsker and Nate Strawn · 2017
Later among the works it cites.
Resilience: A criterion for learning in the presence of arbitrary outliers
Jacob Steinhardt, Moses Charikar, and Gregory Valiant · 2017
Later among the works it cites.
First-order stochastic algorithms for escaping from saddle points in almost linear time
Yi Xu and Tianbao Yang · 2017
Later among the works it cites.
Byzantine stochastic gradient descent
Dan Alistarh, Zeyuan Allen-Zhu, and Jerry Li · 2018
Closest in time.
Asynchronous Byzantine machine learning
Georgios Damaskinos, El Mahdi El Mhamdi, Rachid Guerraoui, Rhicheek Patra, and Mahsa Taziki · 2018
Closest in time.
Sever: A robust meta-algorithm for stochastic optimization
Ilias Diakonikolas, Gautam Kamath, Daniel M Kane, Jerry Li, Jacob Steinhardt, and Alistair Stewart · 2018
Closest in time.
Minimizing nonconvex population risk from rough empirical risk
Chi Jin, Lydia T Liu, Rong Ge, and Michael I Jordan · 2018
Closest in time.
Efficient algorithms for outlier-robust regression
Adam Klivans, Pravesh K Kothari, and Raghu Meka · 2018
Closest in time.
High dimensional robust sparse regression
Liu Liu, Yanyao Shen, Tianyang Li, and Constantine Caramanis · 2018
Closest in time.
Robust estimation via robust gradient estimation
Adarsh Prasad, Arun Sai Suggala, Sivaraman Balakrishnan, and Pradeep Ravikumar · 2018
Closest in time.
Complexity analysis of second-order line-search algorithms for smooth nonconvex optimization
Clément W Royer and Stephen J Wright · 2018
Closest in time.
A newton-cg algorithm with complexity guarantees for smooth unconstrained optimization
Clément W Royer, Michael O’Neill, and Stephen J Wright · 2018
Closest in time.
Securing distributed machine learning in high dimensions
Lili Su and Jiaming Xu · 2018
Closest in time.
Generalized Byzantine-tolerant SGD
Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta · 2018
Closest in time.