Fetching the paper…
Reading the bibliography…
We characterize a prevalent weakness of deep neural networks (DNNs)---overthinking---which occurs when a DNN can reach correct predictions before its final layer.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Thinking, fast and slow
Kahneman, D · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A.-r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Social cognition: From brains to culture
Fiske, S. T. and Taylor, S. E · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Cited alongside, same era.
Deeply-supervised nets
Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Generalizing pooling functions in convolutional neural networks: Mixed, gated, and tree
Lee, C.-Y., Gallagher, P. W., and Tu, Z · 2016
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Later among the works it cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Later among the works it cites.
Multi-scale dense convolutional networks for efficient prediction
Huang, G., Chen, D., Li, T., Wu, F., van der Maaten, L., and Weinberger, K. Q · 2017
Later among the works it cites.
Runtime neural pruning
Lin, J., Rao, Y., Lu, J., and Zhou, J · 2017
Later among the works it cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Branchynet: Fast inference via early exiting from deep neural networks
Teerapittayanon, S., McDanel, B., and Kung, H · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Adaptive neural networks for efficient inference
Bolukbasi, T., Wang, J., Dekel, O., and Saligrama, V · 2017
Cited alongside, same era.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Gu, T., Dolan-Gavitt, B., and Garg, S · 2017
Cited alongside, same era.
Later among the works it cites.
Idk cascades: Fast deep learning by learning not to overthink
Wang, X., Luo, Y., Crankshaw, D., Tumanov, A., Yu, F., and Gonzalez, J. E · 2017
Later among the works it cites.
Greedy layerwise learning can scale to imagenet
Belilovsky, E., Eickenberg, M., and Oyallon, E · 2018
Closest in time.
Learning deep ResNet blocks sequentially using boosting theory
Huang, F., Ash, J., Langford, J., and Schapire, R · 2018
Closest in time.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Papernot, N. and McDaniel, P · 2018
Closest in time.
One bit matters: Understanding adversarial examples as the abuse of redundancy
Wang, J., Jia, R., Friedland, G., Li, B., and Spanos, C · 2018
Closest in time.