Fetching the paper…
Reading the bibliography…
The concept of conditional computation for deep nets has been proposed previously to improve model performance by selectively using only parts of the model conditioned on the sample it is processing.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Principles of neural science
E.R. Kandel, J.H. Schwartz, and T.M. Jessell · 1991
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Adaptive dropout for training deep neural networks
Jimmy Ba and Brendan Frey · 2013
Earlier work this paper cites.
Deep learning of representations: Looking forward
Yoshua Bengio · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Low-rank approximations for conditional feedforward computation in deep neural networks
Andrew Davis and Itamar Arel · 2013
Earlier work this paper cites.
Deep sequential neural network
Ludovic Denoyer and Patrick Gallinari · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Deep networks with internal selective attention through feedback connections
Marijn F Stollenga, Jonathan Masci, Faustino Gomez, and Jürgen Schmidhuber · 2014
Cited alongside, same era.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Łukasz Kaiser and Ilya Sutskever · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, and Vincent Vanhoucke · 2017
Later among the works it cites.
Learning sparse deep feedforward networks via tree skeleton expansion
Zhourong Chen, Xiaopeng Li, and Nevin L. Zhang · 2018
Closest in time.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Can active memory replace attention?
Łukasz Kaiser and Samy Bengio · 2016
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrance Devries and Graham W. Taylor · 2017
Cited alongside, same era.
Shake-shake regularization
Xavier Gastaldi · 2017
Cited alongside, same era.
Łukasz Kaiser and Samy Bengio · 2018
Closest in time.
Fast decoding in sequence models using discrete latent variables
Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, and Noam Shazeer · 2018
Closest in time.
Hydranets: Specialized dynamic architectures for efficient inference
Ravi Teja Mullapudi, William R. Mark, Noam Shazeer, and Kayvon Fatahalian · 2018
Closest in time.
Convolutional networks with adaptive inference graphs
Andreas Veit and Serge J. Belongie · 2018
Closest in time.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le · 2018
Closest in time.