What’s hidden in a randomly weighted neural network?
Ramanujan, V.; Wortsman, M.; Kembhavi, A.; Farhadi, A.; and Rastegari, M. 2020 · 2020
Later among the works it cites.
Green ai
Schwartz, R.; Dodge, J.; Smith, N. A.; and Etzioni, O. 2020 · 2020
Later among the works it cites.
BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning
Wen, Y.; Tran, D.; and Ba, J. 2020 · 2020
Later among the works it cites.
Gans can play lottery tickets too
Original
Chen, X.; Zhang, Z.; Sui, Y.; and Chen, T. 2021 · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Original
Fedus, W.; Zoph, B.; and Shazeer, N. 2021 · 2021
Later among the works it cites.
Highly accurate protein structure prediction with AlphaFold
Jumper, J.; Evans, R.; Pritzel, A.; Green, T.; Figurnov, M.; Ronneberger, O.; Tunyasuvunakool, K.; Bates, R.; Žídek, A.; Potapenko, A.; et al. 2021 · 2021
Later among the works it cites.
Deep ensembling with no overhead for either training or testing: The all-round blessings of dynamic sparsity
Original
Liu, S.; Chen, T.; Atashgahi, Z.; Chen, X.; Sokar, G.; Mocanu, E.; Pechenizkiy, M.; Wang, Z.; and Mocanu, D. C. 2021 · 2021
Later among the works it cites.
Carbon emissions and large neural network training
Original
Patterson, D.; Gonzalez, J.; Le, Q.; Liang, C.; Munguia, L.-M.; Rothchild, D.; So, D.; Texier, M.; and Dean, J. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Original
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
Load-balanced Gather-scatter Patterns for Sparse Deep Neural Networks
Original
Sun, F.; Qin, M.; Zhang, T.; Ma, X.; Li, H.; Luo, J.; Zhao, Z.; Chen, Y.-K.; and Xie, Y. 2021 · 2021
Later among the works it cites.
Mest: Accurate and fast memory-economic sparse training framework on the edge
Yuan, G.; Ma, X.; Niu, W.; Li, Z.; Kong, Z.; Liu, N.; Gong, Y.; Zhan, Z.; He, C.; Jin, Q.; et al. 2021 · 2021
Later among the works it cites.
Coarsening the granularity: Towards structurally sparse lottery tickets
Original
Chen, T.; Chen, X.; Ma, X.; Wang, Y.; and Wang, Z. 2022 · 2022
Closest in time.
Gradient flow in sparse neural networks and how lottery tickets win
Evci, U.; Ioannou, Y.; Keskin, C.; and Dauphin, Y. 2022 · 2022
Closest in time.
Patching open-vocabulary models by interpolating weights
Ilharco, G.; Wortsman, M.; Gadre, S. Y.; Song, S.; Hajishirzi, H.; Kornblith, S.; Farhadi, A.; and Schmidt, L. 2022 · 2022
Closest in time.
The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training
Original
Liu, S.; Chen, T.; Chen, X.; Shen, L.; Mocanu, D. C.; Wang, Z.; and Pechenizkiy, M. 2022 · 2022
Closest in time.
Diverse Weight Averaging for Out-of-Distribution Generalization
Original
Rame, A.; Kirchmeyer, M.; Rahier, T.; Rakotomamonjy, A.; Gallinari, P.; and Cord, M. 2022 · 2022
Closest in time.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M.; Ilharco, G.; Gadre, S. Y.; Roelofs, R.; Gontijo-Lopes, R.; Morcos, A. S.; Namkoong, H.; Farhadi, A.; Carmon, Y.; Kornblith, S.; et al. 2022 · 2022
Closest in time.
Superposing Many Tickets into One: A Performance Booster for Sparse Neural Network Training
Original
Yin, L.; Menkovski, V.; Fang, M.; Huang, T.; Pei, Y.; Pechenizkiy, M.; Mocanu, D. C.; and Liu, S. 2022 · 2022
Closest in time.
A Survey for Efficient Open Domain Question Answering
Original
Zhang, Q.; Chen, S.; Xu, D.; Cao, Q.; Chen, X.; Cohn, T.; and Fang, M. 2022 · 2022
Closest in time.