Fetching the paper…
Reading the bibliography…
Larger networks generally have greater representational power at the cost of increased computational complexity.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
A parallel mixture of svms for very large scale problems
R. Collobert, S. Bengio, and Y. Bengio · 2002
Earlier work this paper cites.
Scaling large learning problems with hard parallel mixtures
R. Collobert, Y. Bengio, and S. Bengio · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus · 2013
Earlier work this paper cites.
K. Cho and Y. Bengio · 2014
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
D. Eigen, M. Ranzato, and I. Sutskever · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Conditional computation in neural networks for faster models
E. Bengio, P.-L. Bacon, J. Pineau, and D. Precup · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Network of experts for large-scale image categorization
K. Ahmed, M. H. Baig, and L. Torresani · 2016
Cited alongside, same era.
On the expressive power of deep learning: A tensor analysis
N. Cohen, O. Sharir, and A. Shashua · 2016
Cited alongside, same era.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Hard mixtures of experts for large scale weakly supervised vision
S. Gross, M. Ranzato, and A. Szlam · 2017
ThiNet: A filter level pruning method for deep neural network compression
J.-H. Luo, J. Wu, and W. Lin · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Later among the works it cites.
Dilated residual networks
F. Yu, V. Koltun, and T. Funkhouser · 2017
Later among the works it cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi · 2018
Closest in time.
Idk cascades: Fast deep learning by learning not to overthink
X. Wang, Y. Luo, D. Crankshaw, A. Tumanov, F. Yu, and J. E. Gonzalez · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Channel pruning for accelerating very deep neural networks
Y. He, X. Zhang, and J. Sun · 2017
Cited alongside, same era.
Data-driven sparse structure selection for deep neural networks
Z. Huang and N. Wang · 2017
Cited alongside, same era.
Pruning filters for efficient convnets
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf · 2017
Cited alongside, same era.
Runtime neural pruning
J. Lin, Y. Rao, J. Lu, and J. Zhou · 2017
Cited alongside, same era.
Closest in time.
Learning to compose topic-aware mixture of experts for zero-shot video captioning
X. Wang, J. Wu, D. Zhang, Y. Su, and W. Y. Wang · 2018
Closest in time.
Skipnet: Learning dynamic routing in convolutional networks
X. Wang, F. Yu, Z.-Y. Dou, T. Darrell, and J. E. Gonzalez · 2018
Closest in time.
Blockdrop: Dynamic inference paths in residual networks
Z. Wu, T. Nagarajan, A. Kumar, S. Rennie, L. S. Davis, K. Grauman, and R. Feris · 2018
Closest in time.
Deep layer aggregation
F. Yu, D. Wang, and T. Darrell · 2018
Closest in time.