Z. Yang and W. U. Bajwa, “BRIDGE: Byzantine-resilient decentralized gradient descent,” arXiv preprint arXiv:1908.08098v1 , Aug. 2019. [Online]. Available: https://arxiv.org/abs/1908.08098v1
Original
1908
Earlier work this paper cites.
J. Verger-Gaugry, “Covering a ball with smaller equal balls in R n R^{n} ,” Discrete & Computational Geometry , vol. 33, no. 1, pp. 143–155, 2005
Original
1908
Earlier work this paper cites.
W. Hoeffding, “Probability inequalities for sums of bounded random variables,” J. American Stat. Assoc. , vol. 58, no. 301, pp. 13–30, 1963
1963
Earlier work this paper cites.
L. Lamport, R. Shostak, and M. Pease, “The Byzantine generals problem,” ACM Trans. Programming Languages and Syst. , vol. 4, no. 3, pp. 382–401, 1982
1982
Earlier work this paper cites.
T. Ypma, “Local convergence of inexact Newton methods,” SIAM Journal on Numerical Analysis , vol. 21, no. 3, pp. 583–590, 1984
1984
Earlier work this paper cites.
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
V. Vapnik, The Nature of Statistical Learning Theory , 2nd ed. New York, NY: Springer-Verlag, 1999
1999
Earlier work this paper cites.
F. Sebastiani, “Machine learning in automated text categorization,” ACM Computing Surveys , vol. 34, no. 1, pp. 1–47, 2002
2002
Earlier work this paper cites.
K. Driscoll, B. Hall, H. Sivencrona, and P. Zumsteq, “Byzantine fault tolerance, from theory to reality,” in Proc. Int. Conf. Computer Safety, Reliability, and Security (SAFECOMP’03) , 2003, pp. 235–248
2003
Earlier work this paper cites.
H. H. Sohrab, Basic Real Analysis , 2nd ed. New York, NY: Springer, 2003
2003
Earlier work this paper cites.
Y. Nesterov, Introductory Lectures on Convex Optimization , ser. Applied optimization; v. 87. Springer US, 2004
2004
Earlier work this paper cites.
P. Dutta, R. Guerraoui, and M. Vukolic, “Best-case complexity of asynchronous Byzantine consensus,” EPFL/IC/200499, Tech. Rep., 2005
2005
Earlier work this paper cites.
J. B. Predd, S. B. Kulkarni, and H. V. Poor, “Distributed learning in wireless sensor networks,” IEEE Signal Process. Mag. , vol. 23, no. 4, pp. 56–69, 2006
2006
Earlier work this paper cites.
S. B. Kotsiantis, I. Zaharakis, and P. Pintelas, “Supervised machine learning: A review of classification techniques,” Emerging Artificial Intell. Applicat. Comput. Eng. , vol. 160, pp. 3–24, 2007
2007
Earlier work this paper cites.
Y. Bengio, “Learning deep architectures for AI,” Found. and Trends Mach. Learning , vol. 2, no. 1, pp. 1–127, 2009
2009
Earlier work this paper cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Trans. Autom. Control , vol. 54, no. 1, pp. 48–61, 2009
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
S. S. Ram, A. Nedić, and V. Veeravalli, “Distributed stochastic subgradient projection algorithms for convex optimization,” J. Optim. Theory and Appl. , vol. 147, no. 3, pp. 516–545, 2010
2010
Earlier work this paper cites.
P. A. Forero, A. Cano, and G. B. Giannakis, “Consensus-based distributed support vector machines,” J. Mach. Learning Research , vol. 11, pp. 1663–1707, 2010
2010
Earlier work this paper cites.
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed optimization and statistical learning via the alternating direction method of multipliers,” Found. and Trends Mach. Learning , vol. 3, no. 1, pp. 1–122, 2011
2011
Earlier work this paper cites.
P. J. Huber, Robust Statistics . Berlin, Heidelberg: Springer, 2011
2011
Earlier work this paper cites.
J. Sousa and A. Bessani, “From Byzantine consensus to BFT state machine replication: A latency-optimal transformation,” in Proc. 9th Euro. Dependable Computing Conf.(EDCC’12) , 2012, pp. 37–48
2012
Earlier work this paper cites.
J. C. Duchi, A. Agarwal, and M. J. Wainwright, “Dual averaging for distributed optimization: Convergence analysis and network scaling,” IEEE Trans. Autom. control , vol. 57, no. 3, pp. 592–606, 2012
2012
Earlier work this paper cites.
J. F. Mota, J. M. Xavier, P. M. Aquiar, and M. Puschel, “D-ADMM: A communication-efficient distributed algorithm for separable optimization,” IEEE Trans. Signal Process. , vol. 61, no. 10, pp. 2718–2723, 2013
2013
Earlier work this paper cites.
H. J. LeBlanc, H. Zhang, X. Koutsoukos, and S. Sundaram, “Resilient asymptotic consensus in robust networks,” IEEE J. Sel. Areas in Commun. , vol. 31, no. 4, pp. 766–781, 2013
2013
Earlier work this paper cites.
A. H. Sayed, “Adaptation, learning, and optimization over networks,” Found. and Trends Mach. Learning , vol. 7, no. 4-5, pp. 311–801, 2014
2014
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in Proc. 11th USENIX Symp. Operating Systems Design and Implementation (OSDI’14) , Broomfield, CO, Oct. 2014, pp. 583–598
2014
Earlier work this paper cites.
W. Shi, Q. Ling, K. Yuan, G. Wu, and W. Yin, “On the linear convergence of the ADMM in decentralized consensus optimization,” IEEE Trans. Signal Process. , vol. 62, no. 7, pp. 1750–1761, 2014
2014
Earlier work this paper cites.
N. H. Vaidya, L. Tseng, and G. Liang, “Iterative Byzantine vector consensus in incomplete graphs,” in Proc. 15th Int. Conf. Distributed Computing and Networking , 2014, pp. 14–28
2014
Earlier work this paper cites.
A. Nedić and A. Olshevsky, “Distributed optimization over time-varying directed graphs,” IEEE Trans. Autom. Control , vol. 60, no. 3, pp. 601–615, 2015
2015
Earlier work this paper cites.
J. Sun, Q. Qu, and J. Wright, “When are nonconvex problems not scary?” arXiv preprint arXiv:1510.06096 , 2015
Original
2015
Earlier work this paper cites.