Fetching the paper…
Reading the bibliography…
Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research.
Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
Shiwei Liu, Decebal Constantin Mocanu, Amarsagar Reddy Ramapuram Matavalam, Yulong Pei, and Mykola Pechenizkiy · 1901
Earlier work this paper cites.
The State of Sparsity in Deep Neural Networks
Trevor Gale, Erich Elsen, and Sara Hooker · 1902
Earlier work this paper cites.
Hesham Mostafa and Xin Wang · 1902
Earlier work this paper cites.
The Generalization-Stability Tradeoff in Neural Network Pruning
Brian R. Bartoldson, Ari S. Morcos, Adrian Barbu, and Gordon Erlebacher · 1906
Earlier work this paper cites.
A Signal Propagation Perspective for Pruning Neural Networks at Initialization
Namhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, and Philip H. S. Torr · 1906
Earlier work this paper cites.
Energy and Policy Considerations for Deep Learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 1906
Earlier work this paper cites.
Sparse Networks from Scratch: Faster Training without Losing Performance
Tim Dettmers and Luke Zettlemoyer · 1907
Earlier work this paper cites.
Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
Chaithanya Kumar Mummadi, Tim Genewein, Dan Zhang, Thomas Brox, and Volker Fischer · 1908
Earlier work this paper cites.
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 1908
Earlier work this paper cites.
DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures
Huanrui Yang, Wei Wen, and Hai Li · 1908
Earlier work this paper cites.
Reducing Transformer Depth on Demand with Structured Dropout
Angela Fan, Edouard Grave, and Armand Joulin · 1909
Earlier work this paper cites.
Structured Pruning of Large Language Models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei · 1910
Earlier work this paper cites.
Rigging the Lottery: Making All Tickets Winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 1911
Earlier work this paper cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin · 1912
Earlier work this paper cites.
One Shot Pruning of Recurrent Neural Networks by Jacobian spectrum evaluation
Matthew Shunshi Zhang and Bradly C Stadie · 1912
Earlier work this paper cites.
Amir Hadifar, Johannes Deleu, Chris Develder, and Thomas Demeester · 2001
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
How fine can fine-tuning be? Learning efficient language models
Evani Radiya-Dixit and Xin Wang · 2004
Earlier work this paper cites.
Movement Pruning: Adaptive Sparsity by Fine-Tuning
Victor Sanh, Thomas Wolf, and Alexander M. Rush · 2005
Earlier work this paper cites.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2006
Cited alongside, same era.
Pruning neural networks without any data by iteratively conserving synaptic flow
Hidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, and Surya Ganguli · 2006
Cited alongside, same era.
Ramanujan Bipartite Graph Products for Efficient Block Sparse Neural Networks
Dharma Teja Vooturi, Girish Varma, and Kishore Kothapalli · 2006
Cited alongside, same era.
The Lottery Ticket Hypothesis for Pre-trained BERT Networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin · 2007
Cited alongside, same era.
Dynamic Channel Pruning: Feature Boosting and Suppression
Xitong Gao, Yiren Zhao, Łukasz Dudziak, Robert Mullins, and Cheng-zhong Xu · 2018
Later among the works it cites.
Measuring the Intrinsic Dimension of Objective Landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Later among the works it cites.
Learning Sparse Neural Networks through L_0 Regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2018
Later among the works it cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle and Michael Carbin · 2019
Later among the works it cites.
Multiplicative Interactions and Where to Find Them
Siddhant M Jayakumar, Wojciech M Czarnecki, Jacob Menick, Jonathan Schwarz, Jack Rae, Simon Osidnero, Yee Whye Teh, Tim Harley, and Razvan Pascanu Deepmind · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin · 2009
Cited alongside, same era.
A Gradient Flow Framework For Analyzing Network Pruning
Ekdeep Singh Lubana and Robert P. Dick · 2009
Cited alongside, same era.
Sanity-Checking Pruning Methods: Random Tickets can Win the Jackpot
Jingtong Su, Yihang Chen, Tianle Cai, Tianhao Wu, Ruiqi Gao, Liwei Wang, and Jason D. Lee · 2009
Cited alongside, same era.
Gradient Flow in Sparse Neural Networks and How Lottery Tickets Win
Utku Evci, Yani A. Ioannou, Cem Keskin, and Yann Dauphin · 2010
Cited alongside, same era.
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta · 2012
Cited alongside, same era.
Deep Rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein · 2017
Cited alongside, same era.
GPU Kernels for Block-Sparse Weights
Scott Gray, Alec Radford, and Diederik P Kingma · 2017
Cited alongside, same era.
Learning Efficient Convolutional Networks through Network Slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang · 2017
Cited alongside, same era.
Later among the works it cites.
SNIP: Single-shot Network Pruning based on Connection Sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr · 2019
Later among the works it cites.
Top-KAST: Top-K Always Sparse Training
Siddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero, and Erich Elsen · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Chaoqi Wang, Guodong Zhang, and Roger Grosse · 2020
Later among the works it cites.
EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, and Jingjing Liu · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Graphcore Homepage, 2021
Graphcore · 2021
Closest in time.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Closest in time.
BASE Layers: Simplifying Training of Large, Sparse Models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Closest in time.
Carbon Emissions and Large Neural Network Training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Closest in time.
AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural Networks
Alexandra Peste, Eugenia Iofinova, Adrian Vladu, and Dan Alistarh · 2021
Closest in time.
Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
Kale-ab Tessera, Sara Hooker, and Benjamin Rosman · 2021
Closest in time.
Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li · 2021
Closest in time.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H. Nguyen, Madeleine Gibescu, and Antonio Liotta · 2041
Closest in time.