Fetching the paper…
Reading the bibliography…
The progress of some AI paradigms such as deep learning is said to be linked to an exponential growth in the number of parameters.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Visualizing and understanding convolutional networks, 2013
M. D. Zeiler and R. Fergus · 2013
Earlier work this paper cites.
Going deeper with convolutions, 2014
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Earlier work this paper cites.
Keras applications, 2015
F. Chollet · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Achieved FLOPs , 2015
C. NVIDIA · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge, 2015
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2015
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision, 2015
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2015
Earlier work this paper cites.
Convnet burden: Estimates of memory consumption and flop counts for various convolutional neural networks., 2016
S. Albanie · 2016
Earlier work this paper cites.
Pretrained models for Pytorch , 2016
R. Cadene · 2016
Earlier work this paper cites.
Evaluating the energy efficiency of deep convolutional neural networks on cpus and gpus
D. Li, X. Chen, M. Becchi, and Z. Zong · 2016
Earlier work this paper cites.
Torchvision models, 2016
A. Paszke, S. Gross, S. Chintala, and G. Chanan · 2016
Earlier work this paper cites.
Inception-v4, inception-resnet and the impact of residual connections on learning, 2016
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi · 2016
Earlier work this paper cites.
An analysis of deep neural network models for practical applications, 2017
A. Canziani, A. Paszke, and E. Culurciello · 2017
Earlier work this paper cites.
Dual path networks, 2017
Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, and J. Feng · 2017
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions, 2017
F. Chollet · 2017
Earlier work this paper cites.
Dawnbench: An end-to-end deep learning benchmark and competition
C. Coleman, D. Narayanan, D. Kang, T. Zhao, J. Zhang, L. Nardi, P. Bailis, K. Olukotun, C. Ré, and M. Zaharia · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. Patwary, M. Ali, Y. Yang, and Y. Zhou · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Earlier work this paper cites.
Nvidia tesla v100 gpu architectur, 2017
C. NVIDIA · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks, 2017
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices, 2017
X. Zhang, X. Zhou, M. Lin, and J. Sun · 2017
Earlier work this paper cites.
Ai and compute
D. Amodei and D. Hernandez · 2018
Earlier work this paper cites.
Benchmark analysis of representative deep neural network architectures
S. Bianco, R. Cadene, L. Celona, and P. Napoletano · 2018
Earlier work this paper cites.
How fast is my model?, 2018
M. Hollemans · 2018
Earlier work this paper cites.
Densely connected convolutional networks, 2018
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger · 2018
Cited alongside, same era.
Progressive neural architecture search, 2018
C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy · 2018
Cited alongside, same era.
Shufflenet v2: Practical guidelines for efficient cnn architecture design, 2018
N. Ma, X. Zhang, H.-T. Zheng, and J. Sun · 2018
Cited alongside, same era.
Exploring the limits of weakly supervised pretraining, 2018
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten · 2018
Cited alongside, same era.
Accounting for the neglected dimensions of ai progress
F. Martinez-Plumed, S. Avin, M. Brundage, A. Dafoe, S. Ó. hÉigeartaigh, and J. Hernández-Orallo · 2018
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations, 2020
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2020
Later among the works it cites.
Openai’s gpt-3 language model: A technical overview
C. Li · 2020
Later among the works it cites.
Mlperf: An industry standard benchmark suite for machine learning performance
P. Mattson, V. J. Reddi, C. Cheng, C. Coleman, G. Diamos, D. Kanter, P. Micikevicius, D. Patterson, G. Schmuelling, H. Tang, et al · 2020
Later among the works it cites.
Turing-nlg: A 17-billion-parameter language model by microsoft, 2020
C. Rosset · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. NVIDIA · 2018
Cited alongside, same era.
Deep learning inference in facebook data centers: Characterization, performance optimizations and hardware implications, 2018
J. Park, M. Naumov, P. Basu, S. Deng, A. Kalaiah, D. Khudia, J. Law, P. Malani, A. Malevich, S. Nadathur, J. Pino, M. Schatz, A. Sidorov, V. Sivakumar, A. Tulloch, X. Wang, Y. Wu, H. Yuen, U. Diril, D. Dzhulgakov, K. Hazelwood, B. Jia, Y. Jia, L. Qiao, V. Rao, N. Rotem, S. Yoo, and M. Smelyanskiy · 2018
Cited alongside, same era.
Deep contextualized word representations, 2018
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Scaling for edge inference of deep neural networks
X. Xu, Y. Ding, S. X. Hu, M. Niemier, J. Cong, Y. Hu, and Y. Shi · 2018
Cited alongside, same era.
Learning transferable architectures for scalable image recognition, 2018
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le · 2018
Cited alongside, same era.
Handbook of deep learning applications , volume 136
V. E. Balas, S. S. Roy, D. Sharma, and P. Samui · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2020
Later among the works it cites.
Flops counter for convolutional networks in pytorch framework , 2020
V. Sovrasov · 2020
Later among the works it cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices, 2020
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou · 2020
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks, 2020
M. Tan and Q. V. Le · 2020
Later among the works it cites.
Reducing machine learning inference cost for pytorch models - aws online tech talks
D. Thomas · 2020
Later among the works it cites.
The computational limits of deep learning
N. C. Thompson, K. Greenewald, K. Lee, and G. F. Manso · 2020
Later among the works it cites.
Fixing the train-test resolution discrepancy: Fixefficientnet, 2020
H. Touvron, A. Vedaldi, M. Douze, and H. Jégou · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification, 2020
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le · 2020
Later among the works it cites.
Bert-of-theseus: Compressing bert by progressive module replacing, 2020
C. Xu, W. Zhou, T. Ge, F. Wei, and M. Zhou · 2020
Later among the works it cites.
Resnest: Split-attention networks, 2020
H. Zhang, C. Wu, Z. Zhang, Y. Zhu, H. Lin, Z. Zhang, Y. Sun, T. He, J. Mueller, R. Manmatha, M. Li, and A. Smola · 2020
Later among the works it cites.
On the opportunities and risks of foundation models, 2021
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Closest in time.
High-performance large-scale image recognition without normalization, 2021
A. Brock, S. De, S. L. Smith, and K. Simonyan · 2021
Closest in time.
Crossvit: Cross-attention multi-scale vision transformer for image classification, 2021
C.-F. Chen, Q. Fan, and R. Panda · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Closest in time.
Greenhouse gas emission intensity of electricity generation in europe
E. E. A. EEA · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows, 2021
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Closest in time.
Meta pseudo labels, 2021
H. Pham, Z. Dai, Q. Xie, M.-T. Luong, and Q. V. Le · 2021
Closest in time.
Bottleneck transformers for visual recognition, 2021
A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani · 2021
Closest in time.
Papers with code imagenet benchmark (image classification), 2021
R. Stojnic and R. Taylor · 2021
Closest in time.
Efficientnetv2: Smaller models and faster training, 2021
M. Tan and Q. V. Le · 2021
Closest in time.
Measuring the occupational impact of ai: tasks, cognitive abilities and ai benchmarks
S. Tolan, A. Pesole, F. Martínez-Plumed, E. Fernández-Macías, J. Hernández-Orallo, and E. Gómez · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet, 2021
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z. Jiang, F. E. Tay, J. Feng, and S. Yan · 2021
Closest in time.
Scaling vision transformers, 2021
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2021
Closest in time.