Fetching the paper…
Reading the bibliography…
We present a method that trains large capacity neural networks with significantly improved accuracy and lower dynamic computational cost.
On the distribution of the two-sample cramér-von mises criterion
T. W. Anderson · 1962
Earlier work this paper cites.
A parallel computation that assigns canonical object-based frames of reference
Geoffrey F. Hinton · 1981
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Unsupervised learning of image transformations
Roland Memisevic and Geoffrey Hinton · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Deep learning of representations: Looking forward
Yoshua Bengio · 2013
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan · 2013
Earlier work this paper cites.
Deep unsupervised network for multimodal perception, representation and classification
Alain Droniou, Serena Ivaldi, and Olivier Sigaud · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Accelerating very deep convolutional networks for classification and detection
Xiangyu Zhang, Jianhua Zou, Kaiming He, and Jian Sun · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Gated networks: an inventory
Olivier Sigaud, Clément Masson, David Filliat, and Freek Stulp · 2016
Cited alongside, same era.
How to train deep variational autoencoders and probabilistic ladder networks
Casper Kaae Sønderby, Tapani Raiko, Lars Maaløe, Søren Kaae Sønderby, and Ole Winther · 2016
Cited alongside, same era.
Branchynet: Fast inference via early exiting from deep neural networks
S. Teerapittayanon, B. McDanel, and H. T. Kung · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean · 2017
Later among the works it cites.
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia · 2017
Later among the works it cites.
Zhourong Chen, Yang Li, Samy Bengio, and Si Si · 2018
Later among the works it cites.
Compressing neural networks using the variational information bottleneck
Bin Dai, Chen Zhu, Baining Guo, and David P. Wipf · 2018
Later among the works it cites.
Dynamic channel pruning: Feature boosting and suppression
Xitong Gao, Yiren Zhao, Lukasz Dudziak, Robert Mullins, and Cheng-Zhong Xu · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille · 2017
Cited alongside, same era.
Spatially adaptive computation time for residual networks
Michael Figurnov, Maxwell D Collins, Yukun Zhu, Li Zhang, Jonathan Huang, Dmitry Vetrov, and Ruslan Salakhutdinov · 2017
Cited alongside, same era.
Channel pruning for accelerating very deep neural networks
Yihui He, Xiangyu Zhang, and Jian Sun · 2017
Cited alongside, same era.
Categorical reparametrization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Cited alongside, same era.
Not all pixels are equal: Difficulty-aware semantic segmentation via deep layer cascade
Xiaoxiao Li, Ziwei Liu, Ping Luo, Chen Change Loy, and Xiaoou Tang · 2017
Cited alongside, same era.
Bayesian compression for deep learning
Christos Louizos, Karen Ullrich, and Max Welling · 2017
Cited alongside, same era.
Later among the works it cites.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Later among the works it cites.
Multi-scale dense networks for resource efficient image classification
Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Weinberger · 2018
Later among the works it cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Later among the works it cites.
Learning sparse neural networks through l0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2018
Later among the works it cites.
Sbnet: Sparse blocks network for fast inference
Mengye Ren, Andrei Pokrovsky, Bin Yang, and Raquel Urtasun · 2018
Later among the works it cites.
Convolutional networks with adaptive inference graphs
Andreas Veit and Serge Belongie · 2018
Later among the works it cites.
Skipnet: Learning dynamic routing in convolutional networks
Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E. Gonzalez · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Closest in time.