Fetching the paper…
Reading the bibliography…
The width of a neural network matters since increasing the width will necessarily increase the model capacity.
Y. Nesterov, “A method for unconstrained convex minimization problem with the rate of convergence o (1/kˆ 2),” in Doklady an ussr , vol. 269, 1 1983, pp. 543–547
1983
Earlier work this paper cites.
A. Krogh and J. A. Hertz, “A simple weight decay can improve generalization,” in NeurIPS , 1991
1991
Earlier work this paper cites.
B. L. Smith and J. T. MacGregor, Collaborative Learning: A Sourcebook for Higher Education . Office of Educational Research and Improvement (ED), Washington, DC: ERIC, 1992, ch. What is collaborative learning
1992
Earlier work this paper cites.
B. E. Rosen, “Ensemble learning using decorrelated neural networks,” Connection science , 1996
1996
Earlier work this paper cites.
N. Ueda and R. Nakano, “Generalization error of ensemble estimators,” in Proceedings of International Conference on Neural Networks , 1996
1996
Earlier work this paper cites.
A. Blum and T. M. Mitchell, “Combining labeled and unlabeled data with co-training,” in COLT . ACM, 1998
1998
Earlier work this paper cites.
Y. Liu and X. Yao, “Simultaneous training of negatively correlated neural networks in an ensemble,” IEEE Trans. Syst. Man Cybern. Part B , 1999
1999
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” University of Toronto, Toronto, Ontario, Tech. Rep. 0, 2009
2009
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NeurIPS , 2012
2012
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in network,” arXiv preprint arXiv:1312.4400 , 2013
2013
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li, “Imagenet large scale visual recognition challenge,” IJCV , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in ICCV , 2015
2015
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” in BMVC , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in ECCV , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in ECCV , 2016
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR . IEEE Computer Society, 2016
2016
Earlier work this paper cites.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, and A. C. Berg, “SSD: single shot multibox detector,” in ECCV , 2016
2016
Earlier work this paper cites.
J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in CVPR , 2016
2016
Earlier work this paper cites.
S. Xie, R. B. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in CVPR , 2017
2017
Cited alongside, same era.
G. Huang, Y. Li, G. Pleiss, Z. Liu, J. E. Hopcroft, and K. Q. Weinberger, “Snapshot ensembles: Train 1, get M for free,” in ICLR , 2017
2017
Cited alongside, same era.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in CVPR , 2017
2017
Cited alongside, same era.
B. Zoph and Q. V. Le, “Neural architecture search with reinforcement learning,” in ICLR , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Tao, “Deep neural network ensembles,” CoRR , 2019
2019
Later among the works it cites.
H. Liu, K. Simonyan, and Y. Yang, “DARTS: differentiable architecture search,” in ICLR , 2019
2019
Later among the works it cites.
J. Yu, L. Yang, N. Xu, J. Yang, and T. S. Huang, “Slimmable neural networks,” in ICLR . OpenReview.net, 2019, pp. 1–12. [Online]. Available: https://openreview.net/forum?id=H1gMCsAqY7
2019
Later among the works it cites.
J. Yu and T. S. Huang, “Universally slimmable networks and improved training techniques,” in ICCV , 2019
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in NeurIPS , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in CVPR , 2017
2017
Cited alongside, same era.
D. Han, J. Kim, and J. Kim, “Deep pyramidal residual networks,” in CVPR , 2017
2017
Cited alongside, same era.
S. Laine and T. Aila, “Temporal ensembling for semi-supervised learning,” in ICLR , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
X. Gastaldi, “Shake-shake regularization,” CoRR , vol. abs/1705.07485, 2017
2017
Cited alongside, same era.
N. Bodla, B. Singh, R. Chellappa, and L. S. Davis, “Soft-nms - improving object detection with one line of code,” in ICCV , 2017
2017
Cited alongside, same era.
2019
Later among the works it cites.
Y. Yamada, M. Iwamura, T. Akiba, and K. Kise, “Shakedrop regularization for deep residual learning,” IEEE Access , 2019
2019
Later among the works it cites.
G. Zhang, C. Wang, B. Xu, and R. B. Grosse, “Three mechanisms of weight decay regularization,” in ICLR , 2019
2019
Later among the works it cites.
A. Golatkar, A. Achille, and S. Soatto, “Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence,” in NeurIPS , 2019
2019
Later among the works it cites.
E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le, “Autoaugment: Learning augmentation strategies from data,” in CVPR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Zhang, Y. N. Dauphin, and T. Ma, “Fixup initialization: Residual learning without normalization,” in ICLR , 2019
2019
Later among the works it cites.
S. Lim, I. Kim, T. Kim, C. Kim, and S. Kim, “Fast autoaugment,” in NeurIPS , 2019
2019
Later among the works it cites.
M. Berman, H. Jégou, V. Andrea, I. Kokkinos, and M. Douze, “MultiGrain: a unified image embedding for classes and instances,” arXiv e-prints , 2019
2019
Later among the works it cites.
E. Real, A. Aggarwal, Y. Huang, and Q. V. Le, “Regularized evolution for image classifier architecture search,” in AAAI , 2019
2019
Later among the works it cites.
2020
Closest in time.
Y. Ge, D. Chen, and H. Li, “Mutual mean-teaching: Pseudo label refinery for unsupervised domain adaptation on person re-identification,” in ICLR , 2020
2020
Closest in time.
T. Yang, S. Zhu, C. Chen, S. Yan, M. Zhang, and A. Willis, “Mutualnet: Adaptive convnet via mutual learning from network width and resolution,” in ECCV , 2020
2020
Closest in time.
2020
Closest in time.
NVIDIA, “Cuda toolkit documentation v11.1.0,” https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#streams , 2020
2020
Closest in time.
Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang, “Random erasing data augmentation,” in AAAI , 2020
2020
Closest in time.
Á. Casado-García and J. Heras, “Ensemble methods for object detection,” in ECAI , 2020
2020
Closest in time.
NVIDIA, “Deep learning examples for tensor cores,” 2020, https://github.com/NVIDIA/DeepLearningExamples/tree/master/PyTorch/Detection/SSD
2020
Closest in time.
R. Solovyev, W. Wang, and T. Gabruseva, “Weighted boxes fusion: Ensembling boxes from different object detection models,” Image and Vision Computing , 2021
2021
Closest in time.