Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Later among the works it cites.
Triple wins: Boosting accuracy, robustness and efficiency together by enabling input-adaptive inference
Original
Ting-Kuei Hu, Tianlong Chen, Haotao Wang, and Zhangyang Wang · 2020
Later among the works it cites.
Top-kast: Top-k always sparse training
Siddhant Jayakumar, Razvan Pascanu, Jack Rae, Simon Osindero, and Erich Elsen · 2020
Later among the works it cites.
Soft threshold weight reparameterization for learnable sparsity
Aditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman, Prateek Jain, Sham Kakade, and Ali Farhadi · 2020
Later among the works it cites.
Finding trainable sparse networks through neural tangent transfer
Tianlin Liu and Friedemann Zenke · 2020
Later among the works it cites.
Towards neural networks that provably know when they don’t know
Alexander Meinke and Matthias Hein · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Comparing rewinding and fine-tuning in neural network pruning
Alex Renda, Jonathan Frankle, and Michael Carbin · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush · 2020
Later among the works it cites.
Sanity-checking pruning methods: Random tickets can win the jackpot
Original
Jingtong Su, Yihang Chen, Tianle Cai, Tianhao Wu, Ruiqi Gao, Liwei Wang, and Jason D Lee · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Original
Hidenori Tanaka, Daniel Kunin, Daniel LK Yamins, and Surya Ganguli · 2020
Later among the works it cites.
Pruning via iterative ranking of sensitivity statistics
Original
Stijn Verdenius, Maarten Stol, and Patrick Forré · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Chaoqi Wang, Guodong Zhang, and Roger Grosse · 2020
Later among the works it cites.
Intriguing properties of adversarial training at scale
Cihang Xie and Alan Yuille · 2020
Later among the works it cites.
Drawing early-bird tickets: Towards more efficient training of deep networks
Haoran You, Chaojian Li, Pengfei Xu, Yonggan Fu, Yue Wang, Xiaohan Chen, Zhangyang Wang, Richard G Baraniuk, and Yingyan Lin · 2020
Later among the works it cites.
Progressive skeletonization: Trimming more fat from a network at initialization
Pau de Jorge, Amartya Sanyal, Harkirat Behl, Philip Torr, Grégory Rogez, and Puneet K. Dokania · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Original
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Pruning neural networks at initialization: Why are we missing the mark?
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2021
Later among the works it cites.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Original
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Later among the works it cites.
Layer-adaptive sparsity for the magnitude-based pruning
Jaeho Lee, Sejun Park, Sangwoo Mo, Sungsoo Ahn, and Jinwoo Shin · 2021
Later among the works it cites.
Keep the gradients flowing: Using gradient flow to study sparse network optimization
Original
Kale-ab Tessera, Sara Hooker, and Benjamin Rosman · 2021
Later among the works it cites.