Fetching the paper…
Reading the bibliography…
This paper introduces EfficientNetV2, a new family of convolutional networks that have faster training speed and better parameter efficiency than previous models.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Collecting a large-scale dataset of fine-grained cars
Krause, J., Deng, J., Stark, M., and Fei-Fei, L · 2013
Earlier work this paper cites.
Rigid-motion scattering for image classification
Sifre, L · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Training imagenet in 3 hours for 25 minutes
Howard, J · 2018
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
Karras, T., Aila, T., Laine, S., and Lehtinen, J · 2018
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and van der Maaten, L · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Earlier work this paper cites.
Mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Earlier work this paper cites.
Learning transferable architectures for scalable image recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V · 2018
Earlier work this paper cites.
Proxylessnas: Direct neural architecture search on target task and hardware
Cai, H., Zhu, L., and Han, S · 2019
Earlier work this paper cites.
Detnas: Neural architecture search on object detection
Chen, Y., Yang, T., Zhang, X., Meng, G., Pan, C., and Sun, J · 2019
Cited alongside, same era.
Neural architecture search: A survey
Elsken, T., Metzen, J. H., and Hutter, F · 2019
Cited alongside, same era.
Efficientnet-edgetpu: Creating accelerator-optimized neural networks with automl
Gupta, S. and Tan, M · 2019
Cited alongside, same era.
Hoffer, E., Weinstein, B., Hubara, I., Ben-Nun, T., Hoefler, T., and Soudry, D · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q. V., and Chen, Z · 2019
Cited alongside, same era.
Tresnet: High performance gpu-dedicated architecture
Ridnik, T., Lawen, H., Noy, A., Baruch, E. B., Sharir, G., and Friedman, I · 2020
Later among the works it cites.
Efficientdet: Scalable and efficient object detection
Tan, M., Pang, R., and Le, Q. V · 2020
Later among the works it cites.
Fixing the train-test resolution discrepancy: Fixefficientnet
Touvron, H., Vedaldi, A., Douze, M., and Jégou, H · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Later among the works it cites.
Mobiledets: Searching for object detection architectures for mobile accelerators
Xiong, Y., Liu, H., Gupta, S., Akin, B., Bender, G., Kindermans, P.-J., Tan, M., Singh, V., and Chen, B · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, C., Chen, L.-C., Schroff, F., Adam, H., Hua, W., Yuille, A., and Fei-Fei, L · 2019
Cited alongside, same era.
Mnasnet: Platform-aware neural architecture search for mobile
Tan, M., Chen, B., Pang, R., Vasudevan, V., and Le, Q. V · 2019
Cited alongside, same era.
Fixing the train-test resolution discrepancy
Touvron, H., Vedaldi, A., Douze, M., and Jégou, H · 2019
Cited alongside, same era.
Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search
Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., Tian, Y., Vajda, P., Jia, Y., and Keutzer, K · 2019
Cited alongside, same era.
Pda: Progressive data augmentation for general robustness of deep neural networks
Yu, H., Liu, A., Liu, X., Li, G., Luo, P., Cheng, R., Yang, J., and Zhang, C · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Cited alongside, same era.
Later among the works it cites.
Resnest: Split-attention networks
Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Lin, H., Zhang, Z., Sun, Y., He, T., Mueller, J., Manmatha, R., Li, M., and Smola, A · 2020
Later among the works it cites.
Lambdanetworks: Modeling long-range interactions without attention
Bello, I · 2021
Closest in time.
Revisiting resnets: Improved training and scaling strategies
Bello, I., Fedus, W., Du, X., Cubuk, E. D., Srinivas, A., Lin, T.-Y., Shlens, J., and Zoph, B · 2021
Closest in time.
High-performance large-scale image recognition without normalization
Brock, A., De, S., Smith, S. L., and Simonyan, K · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Closest in time.
Searching for fast model families on datacenter accelerators
Li, S., Tan, M., Pang, R., Li, A., Cheng, L., Le, Q., and Jouppi, N · 2021
Closest in time.
Shortformer: Better language modeling using shorter inputs
Press, O., Smith, N. A., and Lewis, M · 2021
Closest in time.
Bottleneck transformers for visual recognition
Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A · 2021
Closest in time.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Closest in time.
Pytorch image model
Wightman, R · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Tay, F. E., Feng, J., and Yan, S · 2021
Closest in time.