Fetching the paper…
Reading the bibliography…
In federated learning (FL), weighted aggregation of local models is conducted to generate a global model, and the aggregation weights are normalized (the sum of weights is 1) and proportional to the local data sizes.
Federated learning with matched averaging
Wang, H., Yurochkin, M., Sun, Y., Papailiopoulos, D., and Khazaeni, Y · 2002
Earlier work this paper cites.
On the translocation of masses
Kantorovich, L. V · 2006
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Nocedal, J., Tang, P. T. P., Mudigere, D., and Smelyanskiy, M · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Wilson, A., Podoprikhin, D., Vetrov, D., and Garipov, T · 2018
Earlier work this paper cites.
On the relation between the sharpest directions of dnn loss and the sgd step length
Jastrzębski, S., Kenton, Z., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Gradient diversity: a key ingredient for scalable distributed learning
Yin, D., Pananjady, A., Lam, M., Papailiopoulos, D., Ramchandran, K., and Bartlett, P · 2018
Earlier work this paper cites.
Three mechanisms of weight decay regularization
Zhang, G., Wang, C., Xu, B., and Grosse, R · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Earlier work this paper cites.
Coherent gradients: An approach to understanding generalization in gradient descent-based optimization
Chatterjee, S · 2019
Earlier work this paper cites.
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S · 2019
Cited alongside, same era.
Stiffness: A new perspective on generalization in neural networks
Fort, S., Nowak, P. K., Jastrzebski, S., and Narayanan, S · 2019
Cited alongside, same era.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2019
Cited alongside, same era.
Hyp-rl: Hyperparameter optimization by reinforcement learning
Jomaa, H. S., Grabocka, J., and Schmidt-Thieme, L · 2019
Cited alongside, same era.
Robust federated learning through representation matching and adaptive hyper-parameters
Mostafa, H · 2019
Cited alongside, same era.
Pyhessian: Neural networks through the lens of the hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W · 2020
Later among the works it cites.
Zielinski, P., Krishnan, S., and Chatterjee, S · 2020
Later among the works it cites.
On large-cohort training for federated learning
Charles, Z., Garrett, Z., Huo, Z., Shmulyian, S., and Smith, V · 2021
Later among the works it cites.
Fedbe: Making bayesian model ensemble applicable to federated learning
Chen, H. and Chao, W · 2021
Later among the works it cites.
Efficient sharpness-aware minimization for improved training of neural networks
Du, J., Yan, H., Feng, J., Zhou, J. T., Zhen, L., Goh, R. S. M., and Tan, V. Y · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An improved analysis of training over-parameterized deep neural networks
Zou, D. and Gu, Q · 2019
Cited alongside, same era.
Federated learning based on dynamic regularization
Acar, D. A. E., Zhao, Y., Matas, R., Mattina, M., Whatmough, P., and Saligrama, V · 2020
Cited alongside, same era.
Making coherence out of nothing at all: measuring the evolution of gradient alignment
Chatterjee, S. and Zielinski, P · 2020
Cited alongside, same era.
Distributionally robust federated averaging
Deng, Y., Kamani, M. M., and Mahdavi, M · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2020
Cited alongside, same era.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2020
Cited alongside, same era.
Scaffold: Stochastic controlled averaging for federated learning
Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A. T · 2020
Cited alongside, same era.
Personalized cross-silo federated learning on non-iid data
Huang, Y., Chu, L., Zhou, Z., Wang, L., Liu, J., Pei, J., and Zhang, Y · 2021
Later among the works it cites.
Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Kwon, J., Kim, J., Park, H., and Choi, I. K · 2021
Later among the works it cites.
What can linear interpolation of neural network loss landscapes tell us?
Vlaar, T. and Frankle, J · 2021
Later among the works it cites.
Spherical motion dynamics: Learning dynamics of normalized neural network using sgd and weight decay
Wan, R., Zhu, Z., Zhang, X., and Sun, J · 2021
Later among the works it cites.
A field guide to federated optimization
Wang, J., Charles, Z., Xu, Z., Joshi, G., McMahan, H. B., Al-Shedivat, M., Andrew, G., Avestimehr, S., Daly, K., Data, D., et al · 2021
Later among the works it cites.
Auto-fedavg: learnable federated averaging for multi-institutional medical image segmentation
Xia, Y., Yang, D., Li, W., Myronenko, A., Xu, D., Obinata, H., Mori, H., An, P., Harmon, S., Turkbey, E., et al · 2021
Later among the works it cites.
Critical learning periods in federated learning
Yan, G., Wang, H., and Li, J · 2021
Later among the works it cites.
Improving generalization in federated learning by seeking flat minima
Caldarola, D., Caputo, B., and Ciccone, M · 2022
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2022
Later among the works it cites.
Auto-fedrl: Federated hyperparameter optimization for multi-institutional medical image segmentation
Guo, P., Yang, D., Hatamizadeh, A., Xu, A., Xu, Z., Li, W., Zhao, C., Xu, D., Harmon, S., Turkbey, E., et al · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Lyu, K., Li, Z., and Arora, S · 2022
Later among the works it cites.
Drflm: Distributionally robust federated learning with inter-client noise via local mixup
Wu, B., Liang, Z., Han, Y., Bian, Y., Zhao, P., and Huang, J · 2022
Later among the works it cites.