Fetching the paper…
Reading the bibliography…
Much recent research has been dedicated to improving the efficiency of training and inference for image classification.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Lecture 6.5 - rmsprop
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Lin, M., Chen, Q., and Yan, S · 2013
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, and Darrel, T · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Arpit, D., Zhou, Y., Kota, B. U., and Govindaraju, V · 2016
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
Chollet, F · 2016
Earlier work this paper cites.
Memory-efficient backpropagation through time
Gruslys, A., Munos, R., Danihelka, I., and Graves, A · 2016
Earlier work this paper cites.
SqueezeNet: AlexNet-level accuracy with 50 × \times fewer parameters and < < 0.5MB model size
Iandola, F. N., Han, S., Moskewicz, M. W., Ashraf, K., Dally, W. J., and Keutzer, K · 2016
Earlier work this paper cites.
Deep Roots: Improving CNN efficiency with hierarchical filter groups
Ioannou, Y., Robertson, D., Cipolla, R., and Criminisi, A · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Ulyanov, D., Vedaldi, A., and Lempitsky, V · 2016
Earlier work this paper cites.
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2016
Earlier work this paper cites.
MobileNets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Bo, C., Kalinichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Earlier work this paper cites.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Ioffe, S · 2017
Cited alongside, same era.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elen, E., Garcia, D., Ginsburg, B., Houston, M., Kichaiev, O., Venkatesh, G., and Wu, H · 2017
Cited alongside, same era.
Learning transferable architectures for scalable image recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V · 2017
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Ben-Nun, T. and Hoefler, T · 2018
Cited alongside, same era.
ProxylessNAS: Direct neural architecture search on target task and hardware
Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q. V., and Adam, H · 2019
Later among the works it cites.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M. X., Chen, D., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z · 2019
Later among the works it cites.
Micro-batch training with batch-channel normalization and weight standardization
Qiao, S., Wang, H., Liu, C., Shen, W., and Yuille, A · 2019
Later among the works it cites.
Singh, S. and Krishnan, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cai, H., Zhu, L., and Han, S · 2018
Cited alongside, same era.
AutoAugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2018
Cited alongside, same era.
PipeDream: Fast and efficient pipeline parallel DNN training
Harlap, A., Narayanan, D., Phanishayee, A., Seshadri, V., devanur, N., Ganger, G., and Gibbons, P · 2018
Cited alongside, same era.
Training Imagenet in 3 hours for $25; and CIFAR10 for $0.26
Howard, J · 2018
Cited alongside, same era.
Towards understanding regularization in batch normalization
Luo, P., Wang, X., Shao, W., and Peng, Z · 2018
Cited alongside, same era.
Revisiting small batch training for deep neural networks
Masters, D. and Luschi, C · 2018
Cited alongside, same era.
Regularized evolution for image classifier architecture search
Real, E., Aggarwal, A., Huang, Y., and Le, Q. V · 2018
Cited alongside, same era.
MobileNetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Cited alongside, same era.
Tan, M. and Le, Q. V · 2019
Later among the works it cites.
Fixing the train-test resolution discrepancy
Touvron, H., Vedaldi, A., Douze, M., and Jégou, H · 2019
Later among the works it cites.
UNAS: Differentiable architecture search meets reinforcement learning
Vahdat, A., Mallya, A., Liu, M.-Y., and Kautz, J · 2019
Later among the works it cites.
CutMix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Later among the works it cites.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Later among the works it cites.
Compounding the performance improvements of assembled techniques in a convolutional neural network
Lee, J., Won, T., Lee, T. K., Lee, H., Gu, G., and Hong, K · 2020
Later among the works it cites.
Neural architecture design for GPU-efficient networks
Lin, M., Chen, H., Sun, X., Qian, Q., Li, H., and Jin, R · 2020
Later among the works it cites.
Evolving normalization-activation layers
Liu, H., Brock, A., Simonyan, K., and Le, Q. V · 2020
Later among the works it cites.
Four things everyone should know to improve batch normalization
Summers, C. and Dinneen, M. J · 2020
Later among the works it cites.
Fixing the train-test resolution discrepancy: FixEfficientNet
Touvron, H., Vedaldi, A., Douze, M., and Jégou, H · 2020
Later among the works it cites.
Revisiting resnets: Improved training and scaling strategies
Bello, I., Fedus, W., Du, X., Cubuk, E. D., Srinivas, A., Lin, T.-Y., Shlens, J., and Zoph, B · 2021
Closest in time.
Proxy-normalizing activations to match batch normalization while removing batch dependence
Labatie, A., Masters, D., Eaton-Rosen, Z., and Luschi, C · 2021
Closest in time.
EfficientNetV2: Smaller models and faster training
Tan, M. and Le, Q. V · 2021
Closest in time.
Training data-efficient image transformers & distillation through attention, 2021
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Closest in time.