Fetching the paper…
Reading the bibliography…
Inference in deep neural networks can be computationally expensive, and networks capable of anytime inference are important in mscenarios where the amount of compute or quantity of input data varies over time.
Energy and policy considerations for deep learning in NLP
Strubell, E.; Ganesh, A.; and McCallum, A. 2019 · 1906
Earlier work this paper cites.
Green AI. CoRR abs/1907.10597 (2019)
Schwartz, R.; Dodge, J.; Smith, N. A.; and Etzioni, O. 2019 · 1907
Earlier work this paper cites.
Neural Network Ensembles
Hansen, L.; and Salamon, P. 1990 · 1990
Earlier work this paper cites.
Neural network ensembles, cross validation, and active learning
Krogh, A.; and Vedelsby, J. 1995 · 1995
Earlier work this paper cites.
Bagging predictors
Breiman, L. 1996 · 1996
Earlier work this paper cites.
Generalization error of ensemble estimators
Ueda, N.; and Nakano, R. 1996 · 1996
Earlier work this paper cites.
Optimal ensemble averaging of neural networks
Naftaly, U.; Intrator, N.; and Horn, D. 1997 · 1997
Earlier work this paper cites.
Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
Henderson, P.; Hu, J.; Romoff, J.; Brunskill, E.; Jurafsky, D.; and Pineau, J. 2020 · 2002
Earlier work this paper cites.
Regularized negative correlation learning for neural network ensembles
Chen, H.; and Yao, X. 2009 · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Krizhevsky, A. 2009 · 2009
Earlier work this paper cites.
Ensemble-based classifiers
Rokach, L. 2010 · 2010
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J.; and Caruana, R. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Why M heads are better than one: Training a diverse ensemble of deep networks
Lee, S.; Purushwalkam, S.; Cogswell, M.; Crandall, D.; and Batra, D. 2015 · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Ensembling neural networks: many could be better than all
Zhou, Z.-H.; Wu, J.; and Tang, W. 2002 · 2016
Earlier work this paper cites.
Adaptive neural networks for efficient inference
Bolukbasi, T.; Wang, J.; Dekel, O.; and Saligrama, V. 2017 · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017 · 2017
Earlier work this paper cites.
Snapshot ensembles: Train 1, get m for free
Huang, G.; Li, Y.; Pleiss, G.; Liu, Z.; Hopcroft, J.; and Weinberger, K. 2017 · 2017
Cited alongside, same era.
SplitNet: Learning to semantically split deep networks for parameter reduction and model parallelization
Kim, J.; Park, Y.; Kim, G.; and Hwang, S. J. 2017 · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B.; Pritzel, A.; and Blundell, C. 2017 · 2017
Cited alongside, same era.
Learning efficient convolutional networks through network slimming
Liu, Z.; Li, J.; Shen, Z.; Huang, G.; Yan, S.; and Zhang, C. 2017 · 2017
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I.; and Hutter, F. 2017 · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Xie, S.; Girshick, R.; Dollár, P.; Tu, Z.; and He, K. 2017 · 2017
Searching for mobilenetv3
Howard, A.; Sandler, M.; Chu, G.; Chen, L.-C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. 2019 · 2019
Later among the works it cites.
Improved Techniques for Training Adaptive Deep Networks
Li, H.; Zhang, H.; Qi, X.; Yang, R.; and Huang, G. 2019 · 2019
Later among the works it cites.
Deep tree learning for zero-shot face anti-spoofing
Liu, Y.; Stehouwer, J.; Jourabloo, A.; and Liu, X. 2019 · 2019
Later among the works it cites.
Hydra: an ensemble of convolutional neural networks for geospatial land classification
Minetto, R.; Segundo, M.; and Sarkar, S. 2019 · 2019
Later among the works it cites.
Adaptative Inference Cost With Convolutional Neural Mixture Models
Ruiz, A.; and Verbeek, J. 2019 · 2019
Later among the works it cites.
Patient Knowledge Distillation for BERT Model Compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large scale distributed neural network training through online distillation
Anil, R.; Pereyra, G.; Passos, A.; Ormandi, R.; Dahl, G. E.; and Hinton, G. E. 2018 · 2018
Cited alongside, same era.
Born again neural networks
Furlanello, T.; Lipton, Z. C.; Tschannen, M.; Itti, L.; and Anandkumar, A. 2018 · 2018
Cited alongside, same era.
Uncertainty estimates and multi-hypotheses networks for optical flow
Ilg, E.; Cicek, O.; Galesso, S.; Klein, A.; Makansi, O.; Hutter, F.; and Brox, T. 2018 · 2018
Cited alongside, same era.
Knowledge distillation by on-the-fly native ensemble
Lan, X.; Zhu, X.; and Gong, S. 2018 · 2018
Cited alongside, same era.
ShuffleNet V2: Practical guidelines for efficient CNN architecture design
Ma, N.; Zhang, X.; Zheng, H.-T.; and Sun, J. 2018 · 2018
Cited alongside, same era.
A modern take on the bias-variance tradeoff in neural networks
Neal, B.; Mittal, S.; Baratin, A.; Tantia, V.; Scicluna, M.; Lacoste-Julien, S.; and Mitliagkas, I. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
MnasNet: Platform-aware neural architecture search for mobile
Tan, M.; Chen, B.; Pang, R.; Vasudevan, V.; Sandler, M.; Howard, A.; and Le, Q. 2019 · 2019
Later among the works it cites.
Adaptive neural trees
Tanno, R.; Arulkumaran, K.; Alexander, D.; Criminisi, A.; and Nori, A. 2019 · 2019
Later among the works it cites.
FBNet: Hardware-aware efficient ConvNet design via differentiable neural architecture search
Wu, B.; Dai, X.; Zhang, P.; Wang, Y.; Sun, F.; Wu, Y.; Tian, Y.; Vajda, P.; Jia, Y.; and Keutzer, K. 2019 · 2019
Later among the works it cites.
Universally slimmable networks and improved training techniques
Yu, J.; and Huang, T. 2019 · 2019
Later among the works it cites.
Graph hypernetworks for neural architecture search
Zhang, C.; Ren, M.; and Urtasun, R. 2019 · 2019
Later among the works it cites.
Once for All: Train One Network and Specialize it for Efficient Deployment
Cai, H.; Gan, C.; Wang, T.; Zhang, Z.; and Han, S. 2020 · 2020
Closest in time.
Depth-Adaptive Transformer
Elbayad, M.; Gu, J.; Grave, E.; and Auli, M. 2020 · 2020
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Geiger, M.; Jacot, A.; Spigler, S.; Gabriel, F.; Sagun, L.; d’Ascoli, S.; Biroli, G.; Hongler, C.; and Wyart, M. 2020 · 2020
Closest in time.
Robust Training with Ensemble Consensus
Lee, J.; and Chung, S.-Y. 2020 · 2020
Closest in time.
Ensemble Distribution Distillation
Malinin, A.; Mlodozeniec, B.; and Gales, M. 2020 · 2020
Closest in time.
Tree-CNN: a hierarchical deep convolutional neural network for incremental learning
Roy, D.; Panda, P.; and Roy, K. 2020 · 2020
Closest in time.
Resolution Adaptive Networks for Efficient Inference
Yang, L.; Han, Y.; Chen, X.; Song, S.; Dai, J.; and Huang, G. 2020 · 2020
Closest in time.
Bignas: Scaling up neural architecture search with big single-stage models
Yu, J.; Jin, P.; Liu, H.; Bender, G.; Kindermans, P.-J.; Tan, M.; Huang, T.; Song, X.; Pang, R.; and Le, Q. 2020 · 2020
Closest in time.