Fetching the paper…
Reading the bibliography…
Over-parameterization of deep neural networks (DNNs) has shown high prediction accuracy for many applications.
Exploration and exploitation in evolutionary algorithms: A survey
Matej Črepinšek and et.al · 2013
Earlier work this paper cites.
Ps–abc: A hybrid algorithm based on particle swarm and artificial bee colony for high-dimensional optimization problems
Zhiyong Li and et.al · 2015
Earlier work this paper cites.
The network data repository with interactive graph analytics and visualization
Ryan A. Rossi and Nesreen Ahmed · 2015
Earlier work this paper cites.
Dsd: Dense-sparse-dense training for deep neural networks
Song Han and et.al · 2016
Earlier work this paper cites.
Diverse neural network learns true target functions
Bo Xie and et.al · 2017
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu and et.al · 2018
Earlier work this paper cites.
A systematic dnn weight pruning framework using alternating direction method of multipliers
Tianyun Zhang and et.al · 2018
Earlier work this paper cites.
Deep rewiring: Training very sparse deep networks
Guillaume Bellec and et.al · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee and et.al · 2019
Earlier work this paper cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Hesham Mostafa and et.al · 2019
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and et.al · 2019
Earlier work this paper cites.
Creator governance in social media entertainment
Stuart Cunningham and David Craig · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown and et.al · 2020
Cited alongside, same era.
Picking winning tickets before training by preserving gradient flow
Chaoqi Wang and et.al · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Utku Evci and et.al · 2020
Cited alongside, same era.
Comparing rewinding and fine-tuning in neural network pruning
Alex Renda and et.al · 2020
Cited alongside, same era.
Pruning neural networks without any data by iteratively conserving synaptic flow
Hidenori Tanaka and et.al · 2020
Cited alongside, same era.
Soft threshold weight reparameterization for learnable sparsity
Aditya Kusupati and et.al · 2020
Cited alongside, same era.
Accelerating transformer-based deep learning models on fpgas using column balanced block pruning
Hongwu Peng and et.al · 2021
Later among the works it cites.
Et: re-thinking self-attention for transformer models on gpus
Shiyang Chen and et.al · 2021
Later among the works it cites.
Mest: Accurate and fast memory-economic sparse training framework on the edge
Geng Yuan and et.al · 2021
Later among the works it cites.
Effective model sparsification by scheduled grow-and-prune methods
Xiaolong Ma and et.al · 2021
Later among the works it cites.
Balancing exploration and exploitation with information and randomization
Robert C Wilson and et.al · 2021
Later among the works it cites.
Sparsifying networks via subdifferential inclusion
Sagar Verma and et.al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do we actually need dense over-parameterization? in-time over-parameterization in sparse training
Shiwei Liu and et.al · 2021
Cited alongside, same era.
Accommodating transformer onto fpga: Coupling the balanced model compression and fpga-implementation optimization
Panjie Qi and et.al · 2021
Cited alongside, same era.
Accelerating framework of transformer by hardware design and model compression co-optimization
Panjie Qi and et.al · 2021
Cited alongside, same era.
Co-exploration of graph neural network and network-on-chip design using automl
Daniel Manu and et.al · 2021
Cited alongside, same era.
Carbon emissions and large neural network training
David Patterson and et.al · 2021
Cited alongside, same era.
Prune once for all: Sparse pre-trained language models
Ofir Zafrir and et.al · 2021
Cited alongside, same era.
Sparse training via boosting pruning plasticity with neuroregeneration
Shiwei Liu and et.al · 2021
Later among the works it cites.
Sparse progressive distillation: Resolving overfitting under pretrain-and-finetune paradigm
Shaoyi Huang and et.al · 2022
Closest in time.
A length adaptive algorithm-hardware co-design of transformer on fpga through sparse attention and dynamic pipelining
Hongwu Peng and et.al · 2022
Closest in time.
Towards sparsification of graph neural networks
Hongwu Peng and et.al · 2022
Closest in time.
Sparse double descent: Where network pruning aggravates overfitting
Zheng He and et.al · 2022
Closest in time.