Fetching the paper…
Reading the bibliography…
Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Dalal, N. and Triggs, B · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., Manzagol, P.-A., and Bottou, L · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., and Oliva, A · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A. A · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Mask R-CNN
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., and Belongie, S · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
A short note about kinetics-600
Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A · 2018
Earlier work this paper cites.
AVA: A video dataset of spatio-temporally localized atomic visual actions
Gu, C., Sun, C., Ross, D. A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., et al · 2018
Earlier work this paper cites.
The iNaturalist species classification and detection dataset
Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Cited alongside, same era.
A short note on the kinetics-700 human action dataset
Carreira, J., Noland, E., Hillier, C., and Zisserman, A · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Feichtenhofer, C., Fan, H., Malik, J., and He, K · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Detectron2, 2019
Wu, Y., Kirillov, A., Massa, F., Lo, W.-Y., and Girshick, R · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Cited alongside, same era.
Levit: a vision transformer in convnet’s clothing for faster inference
Graham, B., El-Nouby, A., Touvron, H., Stock, P., Joulin, A., Jégou, H., and Douze, M · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
An end-to-end transformer model for 3d object detection
Misra, I., Girdhar, R., and Joulin, A · 2021
Later among the works it cites.
Vision transformers for dense prediction
Ranftl, R., Bochkovskiy, A., and Koltun, V · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L · 2021
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I · 2020
Cited alongside, same era.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Cited alongside, same era.
RandAugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Cited alongside, same era.
Augment your batch: Improving generalization through instance repetition
Hoffer, E., Ben-Nun, T., Hubara, I., Giladi, N., Hoefler, T., and Soudry, D · 2020
Cited alongside, same era.
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Cited alongside, same era.
Later among the works it cites.
data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M · 2022
Later among the works it cites.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Dong, X., Bao, J., Chen, D., Zhang, W., Yu, N., Yuan, L., Chen, D., and Guo, B · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Feichtenhofer, C., Fan, H., Li, Y., and He, K · 2022
Later among the works it cites.
Mcmae: Masked convolution meets masked autoencoders
Gao, P., Ma, T., Li, H., Lin, Z., Dai, J., and Qiao, Y · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Later among the works it cites.
Milan: Masked image pretraining on language assisted representation
Hou, Z., Sun, F., Chen, Y.-K., Xie, Y., and Kung, S.-Y · 2022
Later among the works it cites.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Later among the works it cites.
Deit iii: Revenge of the vit
Touvron, H., Cord, M., and Jégou, H · 2022
Later among the works it cites.
Masked feature prediction for self-supervised visual pre-training
Wei, C., Fan, H., Xie, S., Wu, C.-Y., Yuille, A., and Feichtenhofer, C · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H · 2022
Later among the works it cites.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., and Xie, S · 2023
Closest in time.