Fetching the paper…
Reading the bibliography…
The performance of deep neural networks is enhanced by ensemble methods, which average the output of several models.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Bagging predictors
L. Breiman · 1996
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Extra: An exact first-order algorithm for decentralized consensus optimization, 2014
W. Shi, Q. Ling, G. Wu, and W. Yin · 2014
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning, 2015
Y. Gal and Z. Ghahramani · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
S. Zhang, A. Choromanska, and Y. LeCun · 2015
Earlier work this paper cites.
Federated optimization: Distributed machine learning for on-device intelligence
J. Konecný, H. B. McMahan, D. Ramage, and P. Richtárik · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2017
Earlier work this paper cites.
Communication-Efficient Learning of Deep Networks from Decentralized Data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y. Arcas · 2017
Earlier work this paper cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
C. Lakshminarayanan and C. Szepesvari · 2018
Earlier work this paper cites.
D 2 : Decentralized training over decentralized data, 2018
H. Tang, X. Lian, M. Yan, C. Zhang, and J. Liu · 2018
Earlier work this paper cites.
Deep ensembles: A loss landscape perspective
S. Fort, H. Hu, and B. Lakshminarayanan · 2019
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization, 2019
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2019
Earlier work this paper cites.
Federated optimization for heterogeneous networks
T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith · 2019
Earlier work this paper cites.
Local sgd converges fast and communicates little, 2019
S. U. Stich · 2019
Cited alongside, same era.
Bayesian nonparametric federated learning of neural networks
M. Yurochkin, M. Agarwal, S. Ghosh, K. Greenewald, N. Hoang, and Y. Khazaeni · 2019
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis, 2020
J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin · 2020
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
T. Lin, S. U. Stich, K. K. Patel, and M. Jaggi · 2020
Cited alongside, same era.
What is being transferred in transfer learning?
B. Neyshabur, H. Sedghi, and C. Zhang · 2020
Cited alongside, same era.
Federated learning with matched averaging, 2020
H. Wang, M. Yurochkin, Y. Sun, D. Papailiopoulos, and Y. Khazaeni · 2020
Cited alongside, same era.
Resist: Layer-wise decomposition of resnets for distributed training
C. Dun, C. R. Wolfe, C. M. Jermaine, and A. Kyrillidis · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models, 2022
M. Li, S. Gururangan, T. Dettmers, M. Lewis, T. Althoff, N. A. Smith, and L. Zettlemoyer · 2022
Later among the works it cites.
Re-basin via implicit sinkhorn differentiation, 2022
F. A. G. Peña, H. R. Medeiros, T. Dubail, M. Aminbeidokhti, E. Granger, and M. Pedersoli · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization, 2022
A. Ramé, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, P. Gallinari, and M. Cord · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Slowmo: Improving communication-efficient distributed sgd with slow momentum
J. Wang, V. Tantia, N. Ballas, and M. Rabbat · 2020
Cited alongside, same era.
Loss surface simplexes for mode connecting volumes and fast ensembling
G. Benton, W. Maddox, S. Lotfi, and A. G. G. Wilson · 2021
Cited alongside, same era.
Cross-gradient aggregation for decentralized learning from non-iid data
Y. Esfandiari, S. Y. Tan, Z. Jiang, A. Balu, E. Herron, C. Hegde, and S. Sarkar · 2021
Cited alongside, same era.
No one representation to rule them all: Overlapping features of training methods
R. Gontijo-Lopes, Y. Dauphin, and E. D. Cubuk · 2021
Cited alongside, same era.
Consensus control for decentralized deep learning
L. Kong, T. Lin, A. Koloskova, M. Jaggi, and S. Stich · 2021
Cited alongside, same era.
Trade-offs of local sgd at scale: An empirical study, 2021
J. J. G. Ortiz, J. Frankle, M. Rabbat, A. Morcos, and N. Ballas · 2021
Cited alongside, same era.
Robust fine-tuning of zero-shot models, 2022
M. Wortsman, G. Ilharco, J. W. Kim, M. Li, S. Kornblith, R. Roelofs, R. Gontijo-Lopes, H. Hajishirzi, A. Farhadi, H. Namkoong, and L. Schmidt · 2022
Later among the works it cites.
Fedsoup: Improving generalization and personalization in federated learning via selective model interpolation, 2023
M. Chen, M. Jiang, Q. Dou, Z. Wang, and X. Li · 2023
Later among the works it cites.
Dart: Diversify-aggregate-repeat training improves generalization of neural networks, 2023
S. Jain, S. Addepalli, P. Sahu, P. Dey, and R. V. Babu · 2023
Later among the works it cites.
Population parameter averaging (papa), 2023
A. Jolicoeur-Martineau, E. Gervais, K. Fatras, Y. Zhang, and S. Lacoste-Julien · 2023
Later among the works it cites.
Repair: Renormalizing permuted activations for interpolation repair, 2023
K. Jordan, H. Sedghi, O. Saukh, R. Entezari, and B. Neyshabur · 2023
Later among the works it cites.
Efficient deep learning: A survey on making deep learning models smaller, faster, and better
G. Menghani · 2023
Later among the works it cites.
A2cid2: Accelerating asynchronous communication in decentralized deep learning
A. Nabli, E. Belilovsky, and E. Oyallon · 2023
Later among the works it cites.
Computation vs. communication scaling for future transformers on future hardware, 2023
S. Pati, S. Aga, M. Islam, N. Jayasena, and M. D. Sinclair · 2023
Later among the works it cites.
Towards a better theoretical understanding of independent subnetwork training, 2023
E. Shulgin and P. Richtárik · 2023
Later among the works it cites.
Model fusion via optimal transport, 2023
S. P. Singh and M. Jaggi · 2023
Later among the works it cites.
Zero++: Extremely efficient collective communication for giant model training, 2023
G. Wang, H. Qin, S. A. Jacobs, C. Holmes, S. Rajbhandari, O. Ruwase, F. Yan, L. Yang, and Y. He · 2023
Later among the works it cites.
Cyclic data parallelism for efficient parallelism of deep neural networks, 2024
L. Fournier and E. Oyallon · 2024
Closest in time.
Harmony in diversity: Merging neural networks with canonical correlation analysis
S. Horoi, A. M. O. Camacho, E. Belilovsky, and G. Wolf · 2024
Closest in time.