Fetching the paper…
Reading the bibliography…
Scalability and privacy are two critical concerns for cross-device federated learning (FL) systems.
Parallel and Distributed Computation: Numerical Methods
D. P. Bertsekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
Twitter sentiment classification using distant supervision
A. Go, R. Bhayani, and L. Huang · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent, 2011
F. Niu, B. Recht, C. Re, and S. J. Wright · 2011
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. B. Brown, D. Song, Ú. Erlingsson, A. Oprea, and C. Raffel · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
The algorithmic foundations of differential privacy
C. Dwork, A. Roth, et al · 2014
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don’t care
S. Chaturapruek, J. C. Duchi, and C. Ré · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Z. Liu, P. Luo, X. Wang, and X. Tang · 2015
Earlier work this paper cites.
On variance reduction in stochastic gradient descent and its asynchronous variants
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola · 2015
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
Practical secure aggregation for federated learning on user-held data
K. A. Bonawitz, V. Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth · 2016
Earlier work this paper cites.
Revisiting distributed synchronous sgd
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Earlier work this paper cites.
Federated learning of deep networks using model averaging
H. B. McMahan, E. Moore, D. Ramage, and B. A. y Arcas · 2016
Earlier work this paper cites.
Prochlo: Strong privacy for analytics in the crowd
A. Bittau, Ú. Erlingsson, P. Maniatis, I. Mironov, A. Raghunathan, D. Lie, M. Rudominer, U. Kode, J. Tinnes, and B. Seefeld · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
S. Jastrzebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. J. Storkey · 2017
Earlier work this paper cites.
ASAGA: Asynchronous Parallel SAGA
R. Leblond, F. Pedregosa, and S. Lacoste-Julien · 2017
Earlier work this paper cites.
Speeding up distributed machine learning using codes
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran · 2017
Earlier work this paper cites.
Perturbed iterate analysis for asynchronous stochastic optimization
H. Mania, X. Pan, D. Papailiopoulos, B. Recht, K. Ramchandran, and M. I. Jordan · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le · 2017
Earlier work this paper cites.
Gradient coding: Avoiding stragglers in distributed learning
R. Tandon, Q. Lei, A. G. Dimakis, and N. Karampatziakis · 2017
Earlier work this paper cites.
Gradient diversity empowers distributed learning
D. Yin, A. Pananjady, M. Lam, D. S. Papailiopoulos, K. Ramchandran, and P. L. Bartlett · 2017
Cited alongside, same era.
Large batch training of convolutional networks
Y. You, I. Gitman, and B. Ginsburg · 2017
Cited alongside, same era.
Asynchronous stochastic gradient descent with delay compensation
S. Zheng, Q. Meng, T. Wang, W. Chen, N. Yu, Z.-M. Ma, and T.-Y. Liu · 2017
Cited alongside, same era.
Leaf: A benchmark for federated settings
S. Caldas, S. M. K. Duddu, P. Wu, T. Li, J. Konečnỳ, H. B. McMahan, V. Smith, and A. Talwalkar · 2018
Cited alongside, same era.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed sgd
S. Dutta, G. Joshi, S. Ghosh, P. Dube, and P. Nagpurkar · 2018
Advances in asynchronous parallel and distributed optimization
M. Assran, A. Aytekin, H. R. Feyzmahdavian, M. Johansson, and M. G. Rabbat · 2020
Later among the works it cites.
Secure single-server aggregation with (poly) logarithmic overhead
J. H. Bell, K. A. Bonawitz, A. Gascón, T. Lepoint, and M. Raykova · 2020
Later among the works it cites.
Z. Chai, Y. Chen, L. Zhao, Y. Cheng, and H. Rangwala · 2020
Later among the works it cites.
Encode, shuffle, analyze privacy revisited: Formalizations and empirical evaluation
Ú. Erlingsson, V. Feldman, I. Mironov, A. Raghunathan, S. Song, K. Talwar, and A. Thakurta · 2020
Later among the works it cites.
Inverting gradients–how easy is it to break privacy in federated learning?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Federated optimization in heterogeneous networks
T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith · 2018
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
X. Lian, W. Zhang, C. Zhang, and J. Liu · 2018
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
T. Lin, S. U. Stich, K. K. Patel, and M. Jaggi · 2018
Cited alongside, same era.
Learning differentially private recurrent language models
H. B. McMahan, D. Ramage, K. Talwar, and L. Zhang · 2018
Cited alongside, same era.
Scaling neural machine translation
M. Ott, S. Edunov, D. Grangier, and M. Auli · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Cited alongside, same era.
Group normalization
Y. Wu and K. He · 2018
Cited alongside, same era.
J. Geiping, H. Bauermeister, H. Dröge, and M. Moeller · 2020
Later among the works it cites.
The non-iid data quagmire of decentralized machine learning
K. Hsieh, A. Phanishayee, O. Mutlu, and P. Gibbons · 2020
Later among the works it cites.
Scaffold: Stochastic controlled averaging for federated learning
S. P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. T. Suresh · 2020
Later among the works it cites.
Cryptonite: A framework for flexible time-series secure aggregation with online fault tolerance
R. Karl, J. Takeshita, and T. Jung · 2020
Later among the works it cites.
On the convergence of fedavg on non-iid data
X. Li, K. Huang, W. Yang, S. Wang, and Z. Zhang · 2020
Later among the works it cites.
Adaptive federated optimization
S. Reddi, Z. Charles, M. Zaheer, Z. Garrett, K. Rush, J. Konečnỳ, S. Kumar, and H. B. McMahan · 2020
Later among the works it cites.
A. Reisizadeh, I. Tziotis, H. Hassani, A. Mokhtari, and R. Pedarsani · 2020
Later among the works it cites.
M. van Dijk, N. V. Nguyen, T. N. Nguyen, L. M. Nguyen, Q. Tran-Dinh, and P. H. Nguyen · 2020
Later among the works it cites.
Is local sgd better than minibatch sgd?
B. Woodworth, K. K. Patel, S. Stich, Z. Dai, B. Bullins, B. Mcmahan, O. Shamir, and N. Srebro · 2020
Later among the works it cites.
Safa: a semi-asynchronous protocol for fast federated learning with low overhead
W. Wu, L. He, W. Lin, R. Mao, C. Maple, and S. A. Jarvis · 2020
Later among the works it cites.
Straggler mitigation in distributed matrix multiplication: Fundamental limits and optimal coding
Q. Yu, M. A. Maddah-Ali, and A. S. Avestimehr · 2020
Later among the works it cites.
Deep leakage from gradients
L. Zhu and S. Han · 2020
Later among the works it cites.
On large-cohort training for federated learning
Z. Charles, Z. Garrett, Z. Huo, S. Shmulyian, and V. Smith · 2021
Closest in time.
Slow and stale gradients can win the race
S. Dutta, J. Wang, and G. Joshi · 2021
Closest in time.
Practical and private (deep) learning without sampling or shuffling
P. Kairouz, B. McMahan, S. Song, O. Thakkar, A. Thakurta, and Z. Xu · 2021
Closest in time.
M. Lam, G.-Y. Wei, D. Brooks, V. J. Reddi, and M. Mitzenmacher · 2021
Closest in time.
Stragglers are not disaster: A hybrid federated learning algorithm with delayed gradients
X. Li, Z. Qu, B. Tang, and Z. Lu · 2021
Closest in time.
Ppfl: privacy-preserving federated learning with trusted execution environments
F. Mo, H. Haddadi, K. Katevas, E. Marin, D. Perino, and N. Kourtellis · 2021
Closest in time.
A field guide to federated optimization
J. Wang, Z. Charles, Z. Xu, G. Joshi, H. B. McMahan, M. Al-Shedivat, G. Andrew, S. Avestimehr, K. Daly, D. Data, et al · 2021
Closest in time.
On the importance of difficulty calibration in membership inference attacks
L. Watson, C. Guo, G. Cormode, and A. Sablayrolles · 2021
Closest in time.