Fetching the paper…
Reading the bibliography…
Overparameterized networks trained to convergence have shown impressive performance in domains such as computer vision and natural language processing.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 1903
Earlier work this paper cites.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 1904
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V Le · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Understanding knowledge distillation in non-autoregressive machine translation
Chunting Zhou, Graham Neubig, and Jiatao Gu · 1911
Earlier work this paper cites.
The expression of a tensor or a polyadic as a sum of products
Frank L Hitchcock · 1927
Earlier work this paper cites.
A class of methods for solving nonlinear simultaneous equations
Charles G Broyden · 1965
Earlier work this paper cites.
Some mathematical notes on three-mode factor analysis
Ledyard R Tucker · 1966
Earlier work this paper cites.
Speech coding based upon vector quantization
Andres Buzo, A Gray, R Gray, and John Markel · 1980
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Kunihiko Fukushima · 1980
Earlier work this paper cites.
Linear systems , volume 156
Thomas Kailath · 1980
Earlier work this paper cites.
Least squares quantization in pcm
Stuart Lloyd · 1982
Earlier work this paper cites.
Data compression using adaptive coding and partial string matching
John Cleary and Ian Witten · 1984
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Michael C Mozer and Paul Smolensky · 1989
Earlier work this paper cites.
The cascade-correlation learning architecture
Scott E Fahlman and Christian Lebiere · 1990
Earlier work this paper cites.
A simple procedure for pruning back-propagation trained neural networks
Ehud D Karnin · 1990
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla · 1990
Earlier work this paper cites.
Genetic algorithms and neural networks: Optimizing connections and connectivity
Darrell Whitley, Timothy Starkweather, and Christopher Bogart · 1990
Earlier work this paper cites.
A comparative analysis of selection schemes used in genetic algorithms
David E Goldberg and Kalyanmoy Deb · 1991
Earlier work this paper cites.
Generalization by weight-elimination with application to forecasting
Andreas S Weigend, David E Rumelhart, and Bernardo A Huberman · 1991
Earlier work this paper cites.
Simplifying neural networks by soft weight-sharing
Steven J Nowlan and Geoffrey E Hinton · 1992
Earlier work this paper cites.
Removal of hidden units and weights for back propagation networks
Masafumi Hagiwara · 1993
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G Stork · 1993
Earlier work this paper cites.
Pruning algorithms-a survey
Russell Reed · 1993
Earlier work this paper cites.
Optimal brain surgeon: Extensions and performance comparisons
Babak Hassibi, David G Stork, and Gregory Wolff · 1994
Earlier work this paper cites.
Fast pruning using principal components
Asriel U Levin, Todd K Leen, and John E Moody · 1994
Earlier work this paper cites.
Particle swarm optimization
James Kennedy and Russell Eberhart · 1995
Earlier work this paper cites.
An iterative pruning algorithm for feedforward neural networks
Giovanna Castellano, Anna Maria Fanelli, and Marcello Pelillo · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Pruned neural networks for regression
Rudy Setiono and Wee Kheng Leow · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
A new pruning heuristic based on variance analysis of sensitivity information
Andries Petrus Engelbrecht · 2001
Earlier work this paper cites.
Pruning neural networks with distribution estimation algorithms
Erick Cantu-Paz · 2003
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
A node pruning algorithm based on a fourier amplitude sensitivity test method
Philippe Lauret, Eric Fock, and Thierry Alex Mara · 2006
Earlier work this paper cites.
Decompositions of a higher-order tensor in block terms—part ii: Definitions and uniqueness
Lieven De Lathauwer · 2008
Earlier work this paper cites.
An integrated growing-pruning method for feedforward network training
Pramod L Narasimha, Walter H Delashmit, Michael T Manry, Jiang Li, and Francisco Maldonado · 2008
Earlier work this paper cites.
Hash kernels for structured data
Qinfeng Shi, James Petterson, Gideon Dror, John Langford, Alex Smola, and SVN Vishwanathan · 2009
Earlier work this paper cites.
Enhancing the generalization ability of neural networks through controlling the hidden layers
Weishui Wan, Shingo Mabu, Kaoru Shimada, Kotaro Hirasawa, and Jinglu Hu · 2009
Earlier work this paper cites.
Feature hashing for large scale multitask learning
Kilian Weinberger, Anirban Dasgupta, John Langford, Alex Smola, and Josh Attenberg · 2009
Earlier work this paper cites.
The horseshoe estimator for sparse signals
Carlos M Carvalho, Nicholas G Polson, and James G Scott · 2010
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2010
Earlier work this paper cites.
A neural network pruning method optimized with pso algorithm
Juanjuan Tu, Yongzhao Zhan, and Fei Han · 2010
Earlier work this paper cites.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Nathan Halko, Per-Gunnar Martinsson, and Joel A Tropp · 2011
Earlier work this paper cites.
Tensor-train decomposition
Ivan V Oseledets · 2011
Earlier work this paper cites.
Multiresolution mixture modeling using merging of mixture components
Prem Raj Adhikari and Jaakko Hollmen · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Leonard, and Aaron Courville · 2013
Earlier work this paper cites.
Understanding deep architectures using a recursive convolutional network
David Eigen, Jason Rolfe, Rob Fergus, and Yann LeCun · 2013
Earlier work this paper cites.
Multi-digit number recognition from street view imagery using deep convolutional neural networks
Ian J Goodfellow, Yaroslav Bulatov, Julian Ibarz, Sacha Arnoud, and Vinay Shet · 2013
Earlier work this paper cites.
A structure optimisation algorithm for feedforward neural network construction
Hong-Gui Han and Jun-Fei Qiao · 2013
Earlier work this paper cites.
Learning separable filters
Roberto Rigamonti, Amos Sironi, Vincent Lepetit, and Pascal Fua · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran · 2013
Earlier work this paper cites.
Peter huttenlocher (1931–2013), 2013
Christopher A Walsh · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Jian Xue, Jinyu Li, and Yifan Gong · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2014
Earlier work this paper cites.
Convolutional neural network architectures for matching natural language sentences
Baotian Hu, Zhengdong Lu, Hang Li, and Qingcai Chen · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Singular value decomposition based low-footprint speaker adaptation and personalization for deep neural network
Jian Xue, Jinyu Li, Dong Yu, Mike Seltzer, and Yifan Gong · 2014
Earlier work this paper cites.
Compressing neural networks with the hashing trick
Wenlin Chen, James Wilson, Stephen Tyree, Kilian Weinberger, and Yixin Chen · 2015
Earlier work this paper cites.
High-performance hardware for machine learning
William Dally · 2015
Earlier work this paper cites.
8-bit approximations for parallelism in deep learning
Tim Dettmers · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Diederik P Kingma, Tim Salimans, and Max Welling · 2015
Cited alongside, same era.
Tensorizing neural networks
Alexander Novikov, Dmitrii Podoprikhin, Anton Osokin, and Dmitry P Vetrov · 2015
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
Accelerating convolutional networks via global & dynamic filter pruning
Shaohui Lin, Rongrong Ji, Yuchao Li, Yongjian Wu, Feiyue Huang, and Baochang Zhang · 2018
Later among the works it cites.
Neural architecture optimization
Renqian Luo, Fei Tian, Tao Qin, Enhong Chen, and Tie-Yan Liu · 2018
Later among the works it cites.
Piggyback: Adding multiple tasks to a single, fixed network by learning to mask
Arun Mallya and Svetlana Lazebnik · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Cited alongside, same era.
Distilling knowledge from ensembles of neural networks for speech recognition
Yevgen Chebotar and Austin Waters · 2016
Cited alongside, same era.
Loss-aware binarization of deep networks
Lu Hou, Quanming Yao, and James T Kwok · 2016
Cited alongside, same era.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2016
Cited alongside, same era.
Deeply-recursive convolutional network for image super-resolution
Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush · 2016
Cited alongside, same era.
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Later among the works it cites.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Later among the works it cites.
Model compression via distillation and quantization
Antonio Polino, Razvan Pascanu, and Dan Alistarh · 2018
Later among the works it cites.
Neural compatibility modeling with attentive knowledge distillation
Xuemeng Song, Fuli Feng, Xianjing Han, Xin Yang, Wei Liu, and Liqiang Nie · 2018
Later among the works it cites.
Faster gaze prediction with dense networks and fisher pruning
Lucas Theis, Iryna Korshunova, Alykhan Tejani, and Ferenc Huszar · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Colin Wei, Jason Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
Releq: An automatic reinforcement learning approach for deep quantization of neural networks
Amir Yazdanbakhsh, Ahmed T Elthakeb, Prannoy Pilligundla, FatemehSadat Mireshghallah, and Hadi Esmaeilzadeh · 2018
Later among the works it cites.
Learning compact recurrent neural networks with block-term tensor decomposition
Jinmian Ye, Linnan Wang, Guangxi Li, Di Chen, Shandian Zhe, Xinqi Chu, and Zenglin Xu · 2018
Later among the works it cites.
Nisp: Pruning networks using neuron importance score propagation
Ruichi Yu, Ang Li, Chun-Fu Chen, Jui-Hsin Lai, Vlad I Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, and Larry S Davis · 2018
Later among the works it cites.
Explicit loss-error-aware quantization for low-bit deep neural networks
Aojun Zhou, Anbang Yao, Kuan Wang, and Yurong Chen · 2018
Later among the works it cites.
Knowledge distillation by on-the-fly native ensemble
Xiatian Zhu, Shaogang Gong, et al · 2018
Later among the works it cites.
Compressing gans using knowledge distillation
Angeline Aguinaldo, Ping-Yeh Chiang, Alex Gain, Ameya Patil, Kolten Pearson, and Soheil Feizi · 2019
Later among the works it cites.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Later among the works it cites.
Distilling the knowledge of bert for text generation
Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu, and Jingjing Liu · 2019
Later among the works it cites.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Later among the works it cites.
Recurrent stacking of layers for compact neural machine translation models
Raj Dabre and Atsushi Fujita · 2019
Later among the works it cites.
Nest: A neural network synthesis tool based on a grow-and-prune paradigm
Xiaoliang Dai, Hongxu Yin, and Niraj K Jha · 2019
Later among the works it cites.
Exact expressions for double descent and implicit regularization via surrogate random design
Michal Derezinski, Feynman Liang, and Michael W Mahoney · 2019
Later among the works it cites.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and Luke Zettlemoyer · 2019
Later among the works it cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2019
Later among the works it cites.
Fast sparse convnets, 2019
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan · 2019
Later among the works it cites.
The lottery ticket hypothesis at scale
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin · 2019
Later among the works it cites.
Weight agnostic neural networks
Adam Gaier and David Ha · 2019
Later among the works it cites.
Differentiable soft quantization: Bridging full-precision and low-bit neural networks
Ruihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li, Peng Hu, Jiazhen Lin, Fengwei Yu, and Junjie Yan · 2019
Later among the works it cites.
Variational student: Learning compact and sparser networks in knowledge distillation framework
Srinidhi Hegde, Ranjitha Prasad, Ramya Hebbalaguppe, and Vishwajith Kumar · 2019
Later among the works it cites.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Byeongho Heo, Minsik Lee, Sangdoo Yun, and Jin Young Choi · 2019
Later among the works it cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Later among the works it cites.
A study of bfloat16 for deep learning training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, et al · 2019
Later among the works it cites.
Convolutional neural networks with layer reuse
Okan Köpüklü, Maryam Babaee, Stefan Hörmann, and Gerhard Rigoll · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Later among the works it cites.
Towards optimal structured cnn pruning via generative adversarial learning
Shaohui Lin, Rongrong Ji, Chenqian Yan, Baochang Zhang, Liujuan Cao, Qixiang Ye, Feiyue Huang, and David Doermann · 2019
Later among the works it cites.
Mixed precision training with 8-bit floating point
Naveen Mellempudi, Sudarshan Srinivasan, Dipankar Das, and Bharat Kaul · 2019
Later among the works it cites.
Improved knowledge distillation via teacher assistant: Bridging the gap between student and teacher
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, and Hassan Ghasemzadeh · 2019
Later among the works it cites.
Hesham Mostafa and Xin Wang · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Muller, Simon Kornblith, and Geoffrey E Hinton · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Later among the works it cites.
Asap: Architecture search, anneal and prune
Asaf Noy, Niv Nayman, Tal Ridnik, Nadav Zamir, Sivan Doveh, Itamar Friedman, Raja Giryes, and Lihi Zelnik-Manor · 2019
Later among the works it cites.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Later among the works it cites.
Zero: Memory optimization towards training a trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2019
Later among the works it cites.
Turing-nlg: A 17-billion-parameter language model by microsoft
C Rosset · 2019
Later among the works it cites.
Smaller, faster, cheaper, lighter: Introducing distilbert, a distilled version of bert, 2019
Victor Sanh · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Later among the works it cites.
Learning implicitly recurrent cnns through parameter sharing
Pedro Savarese and Michael Maire · 2019
Later among the works it cites.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox · 2019
Later among the works it cites.
Q-bert: Hessian based ultra low precision quantization of bert
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2019
Later among the works it cites.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
And the bit goes down: Revisiting the quantization of neural networks
Pierre Stock, Armand Joulin, Remi Gribonval, Benjamin Graham, and Herve Jegou · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Later among the works it cites.
Efficientnet: Improving accuracy and efficiency through automl and model scaling, 2019a
Mingxing Tan and Quoc V. Le · 2019
Later among the works it cites.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
Similarity-preserving knowledge distillation
Frederick Tung and Greg Mori · 2019
Later among the works it cites.
Sharing attention weights for fast transformer
Tong Xiao, Yinqiao Li, Jingbo Zhu, Zhengtao Yu, and Tongran Liu · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Drawing early-bird tickets: Towards more efficient training of deep networks
Haoran You, Chaojian Li, Pengfei Xu, Yonggan Fu, Yue Wang, Xiaohan Chen, Yingyan Lin, Zhangyang Wang, and Richard G Baraniuk · 2019
Later among the works it cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Closest in time.
Shifted and squeezed 8-bit floating point format for low-precision training of deep neural networks
Leopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Mehran Nekuii, Oguz H Elibol, and Hanlin Tang · 2020
Closest in time.
Big self-supervised models are strong semi-supervised learners, 2020
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Training with quantization noise for extreme model compression, 2020
Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Remi Gribonval, Herve Jegou, and Armand Joulin · 2020
Closest in time.
Compressing deep neural networks via layer fusion, 2020
James O’ Neill, Greg Ver Steeg, and Aram Galstyan · 2020
Closest in time.
Improved noisy student training for automatic speech recognition, 2020
Daniel S. Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V. Le · 2020
Closest in time.
Shapeshifter networks: Cross-layer parameter sharing for scalable and effective deep learning
Bryan A Plummer, Nikoli Dryden, Julius Frost, Torsten Hoefler, and Kate Saenko · 2020
Closest in time.
What’s hidden in a randomly weighted neural network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari · 2020
Closest in time.
Pruning neural networks without any data by iteratively conserving synaptic flow
Hidenori Tanaka, Daniel Kunin, Daniel LK Yamins, and Surya Ganguli · 2020
Closest in time.