Fetching the paper…
Reading the bibliography…
In established network architectures, shortcut connections are often used to take the outputs of earlier layers as additional inputs to later layers.
L. R. Dice, “Measures of the amount of ecologic association between species,” Ecology
1945
Earlier work this paper cites.
D. Harrison Jr and D. L. Rubinfeld, “Hedonic housing prices and the demand for clean air,” Journal of environmental economics and management
1978
Earlier work this paper cites.
K.-I. Funahashi, “On the approximate realization of continuous mappings by neural networks,” Neural networks
1989
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural networks
1989
Earlier work this paper cites.
V. Tikhomirov, “On the representation of continuous functions of several variables as superpositions of continuous functions of one variable and addition,” in Selected Works of AN Kolmogorov
1991
Earlier work this paper cites.
B. Hamann and J.-L. Chen, “Data point selection for piecewise linear curve approximation,” Computer Aided Geometric Design
1994
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition
2009
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” IEEE Transactions on audio, speech, and language processing
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems
2012
Earlier work this paper cites.
V. Vovk, “Kernel ridge regression,” in Empirical inference
2013
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in network,” International Conference on Learning Representations
2014
Earlier work this paper cites.
L. Szymanski and B. McCane, “Deep networks are effective encoders of periodicity,” IEEE transactions on neural networks and learning systems
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” International Conference on Learning Representations
2015
Earlier work this paper cites.
B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik, “Hypercolumns for object segmentation and fine-grained localization,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2015
Earlier work this paper cites.
R. K. Srivastava, K. Greff, and J. Schmidhuber, “Training very deep networks,” in Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 2
2015
Earlier work this paper cites.
B. Neyshabur, R. Tomioka, and N. Srebro, “Norm-based capacity control in neural networks,” in Conference on Learning Theory
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International conference on machine learning
2015
Earlier work this paper cites.
A. Kumar, O. Irsoy, P. Ondruska, M. Iyyer, J. Bradbury, I. Gulrajani, V. Zhong, R. Paulus, and R. Socher, “Ask me anything: Dynamic memory networks for natural language processing,” in International conference on machine learning
2016
Earlier work this paper cites.
G. Wang, “A perspective on deep imaging,” Ieee Access
2016
Earlier work this paper cites.
M. Anthimopoulos, S. Christodoulidis, L. Ebner, A. Christe, and S. Mougiakakou, “Lung pattern classification for interstitial lung diseases using a deep convolutional neural network,” IEEE transactions on medical imaging
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2016
Earlier work this paper cites.
H. N. Mhaskar and T. Poggio, “Deep vs. shallow networks: An approximation theory perspective,” Analysis and Applications
2016
Earlier work this paper cites.
R. Eldan and O. Shamir, “The power of depth for feedforward neural networks,” in Conference on learning theory
2016
Earlier work this paper cites.
A. Veit, M. Wilber, and S. Belongie, “Residual networks behave like ensembles of relatively shallow networks,” in Proceedings of the 30th International Conference on Neural Information Processing Systems
2016
Earlier work this paper cites.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in European conference on computer vision
2016
Earlier work this paper cites.
H. Chen, Y. Zhang, M. K. Kalra, F. Lin, Y. Chen, P. Liao, J. Zhou, and G. Wang, “Low-dose ct with a residual encoder-decoder convolutional neural network,” IEEE transactions on medical imaging
2017
Cited alongside, same era.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2017
Cited alongside, same era.
V. Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE transactions on pattern analysis and machine intelligence
2017
Cited alongside, same era.
G. Larsson, M. Maire, and G. Shakhnarovich, “Fractalnet: Ultra-deep neural networks without residuals,” International Conference on Learning Representations
2017
Cited alongside, same era.
E. Kang, H. J. Koo, D. H. Yang, J. B. Seo, and J. C. Ye, “Cycle-consistent adversarial denoising network for multiphase coronary ct angiography,” Medical physics
2019
Closest in time.
C. You, G. Li, Y. Zhang, X. Zhang, H. Shan, M. Li, S. Ju, Z. Zhao, Z. Zhang, W. Cong, et al
2019
Closest in time.
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever, “Deep double descent: Where bigger models and more data hurt,” in International Conference on Learning Representations
2019
Closest in time.
S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, and R. Wang, “On exact computation with an infinitely wide neural net,” in Advances in Neural Information Processing Systems
2019
Closest in time.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International Conference on Machine Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Liang and R. Srikant, “Why deep neural networks for function approximation?,” in International Conference on Learning Representations
2017
Cited alongside, same era.
Z. Lu, H. Pu, F. Wang, Z. Hu, and L. Wang, “The expressive power of neural networks: a view from the width,” in Proceedings of the 31st International Conference on Neural Information Processing Systems
2017
Cited alongside, same era.
P. L. Bartlett, D. J. Foster, and M. Telgarsky, “Spectrally-normalized margin bounds for neural networks,” in Proceedings of the 31st International Conference on Neural Information Processing Systems
2017
Cited alongside, same era.
D. Smilkov, N. Thorat, B. Kim, F. Viégas, and M. Wattenberg, “Smoothgrad: removing noise by adding noise,” International conference on machine learning
2017
Cited alongside, same era.
M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” International conference on machine learning
2017
Cited alongside, same era.
F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size,” International Conference on Learning Representations
2017
Cited alongside, same era.
D. Rolnick and M. Tegmark, “The power of deeper networks for expressing natural functions,” International Conference on Learning Representations
2018
Cited alongside, same era.
H. Lin and S. Jegelka, “Resnet with one-neuron hidden layers is a universal approximator,” Advances in Neural Information Processing Systems
2018
Cited alongside, same era.
2019
Closest in time.
Y. Chen, H. Fan, B. Xu, Z. Yan, Y. Kalantidis, M. Rohrbach, S. Yan, and J. Feng, “Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution,” in Proceedings of the IEEE International Conference on Computer Vision
2019
Closest in time.
Y. Li, Z. Kuang, Y. Chen, and W. Zhang, “Data-driven neuron allocation for scale aggregation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2019
Closest in time.
S. Xie, A. Kirillov, R. Girshick, and K. He, “Exploring randomly wired neural networks for image recognition,” in Proceedings of the IEEE International Conference on Computer Vision
2019
Closest in time.
E. Real, A. Aggarwal, Y. Huang, and Q. V. Le, “Regularized evolution for image classifier architecture search,” in Proceedings of the aaai conference on artificial intelligence
2019
Closest in time.
B. Wu, X. Dai, P. Zhang, Y. Wang, F. Sun, Y. Wu, Y. Tian, P. Vajda, Y. Jia, and K. Keutzer, “Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2019
Closest in time.
G. Wang, K. Wang, and L. Lin, “Adaptively connected neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2019
Closest in time.
S. Srinivas and F. Fleuret, “Full-gradient representation for neural network visualization,” in Advances in Neural Information Processing Systems
2019
Closest in time.
F. Fan, J. Xiong, and G. Wang, “Universal approximation with quadratic deep networks,” Neural Networks
2020
Closest in time.
F. He, T. Liu, and D. Tao, “Why resnet works? residuals generalize.,” IEEE transactions on neural networks and learning systems
2020
Closest in time.
Z. Huang, S. Liang, M. Liang, and H. Yang, “Dianet: Dense-and-implicit attention network.,” in AAAI
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
I. Bello, “Lambdanetworks: Modeling long-range interactions without attention,” in International Conference on Learning Representations
2020
Closest in time.
K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, and C. Xu, “Ghostnet: More features from cheap operations,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
Closest in time.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
Closest in time.
J. Fang, Y. Sun, Q. Zhang, Y. Li, W. Liu, and X. Wang, “Densely connected search space for more flexible neural architecture search,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
Closest in time.
Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, “Eca-net: Efficient channel attention for deep convolutional neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
Closest in time.
2020
Closest in time.
B.-J. Hou and Z.-H. Zhou, “Learning with interpretable structure from gated rnn,” IEEE transactions on neural networks and learning systems
2020
Closest in time.
A. B. Arrieta, N. Díaz-Rodríguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. García, S. Gil-López, D. Molina, R. Benjamins, et al
2020
Closest in time.
J. Schmidt-Hieber, “The kolmogorov–arnold representation theorem revisited,” Neural Networks
2021
Closest in time.
under review
I. Bello, “Lambdanetworks: Modeling long-range interactions without attention,” in Submitted to International Conference on Learning Representations · 2021
Closest in time.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning
2021
Closest in time.