Fetching the paper…
Reading the bibliography…
Introduced by Hinton et al.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R · 1996
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C · 2012
Earlier work this paper cites.
Adaptive dropout for training deep neural networks
Ba, J. and Frey, B · 2013
Earlier work this paper cites.
Understanding dropout
Baldi, P. and Sadowski, P. J · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., Cun, Y. L., and Fergus, R · 2013
Earlier work this paper cites.
Fast dropout training
Wang, S. and Manning, C · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Gool, L. V · 2014
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2014
Earlier work this paper cites.
Annealed dropout training of deep networks
Rennie, S. J., Goel, V., and Thomas, S · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Efficient object localization using convolutional networks
Tompson, J., Goroshin, R., Jain, A., LeCun, Y., and Bregler, C · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Fractalnet: Ultra-deep neural networks without residuals
Larsson, G., Maire, M., and Shakhnarovich, G · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
DeVries, T. and Taylor, G. W · 2017
Cited alongside, same era.
Concrete dropout
Gal, Y., Hron, J., and Kendall, A · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Mask R-CNN
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Cited alongside, same era.
Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization
Pytorch image models
Wightman, R · 2019
Later among the works it cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Later among the works it cites.
Lookahead optimizer: k steps forward, 1 step back
Zhang, M., Lucas, J., Ba, J., and Hinton, G. E · 2019
Later among the works it cites.
Semantic understanding of scenes through the ADE20K dataset
Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., and Torralba, A · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kumar Singh, K. and Jae Lee, Y · 2017
Cited alongside, same era.
Learning efficient convolutional networks through network slimming
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., and Zhang, C · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Cited alongside, same era.
Curriculum dropout
Morerio, P., Cavazza, J., Volpi, R., Vidal, R., and Murino, V · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2020
Later among the works it cites.
MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark
MMSegmentation-contributors · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Later among the works it cites.
Random erasing data augmentation
Zhong, Z., Zheng, L., Kang, G., Li, S., and Yang, Y · 2020
Later among the works it cites.
An empirical study of training self-supervised Vision Transformers
Chen, X., Xie, S., and He, K · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L · 2021
Later among the works it cites.
Efficientnetv2: Smaller models and faster training
Tan, M. and Le, Q · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Later among the works it cites.
Going deeper with image transformers
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and Jégou, H · 2021
Later among the works it cites.
When vision transformers outperform resnets without pre-training or strong data augmentations
Chen, X., Hsieh, C.-J., and Gong, B · 2022
Later among the works it cites.
Adaptive stochastic variance reduction for non-convex finite-sum minimization
Kavis, A., Skoulakis, S., Antonakopoulos, K., Dadi, L. T., and Cevher, V · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Mu, N., Kirillov, A., Wagner, D., and Xie, S · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.