Fetching the paper…
Reading the bibliography…
Masked image modeling (MIM) as pre-training is shown to be effective for numerous vision downstream tasks, but how and where MIM works remain unclear.
What does bert look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D. (2019) · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Visualizing and understanding the effectiveness of bert
Hao, Y., Dong, L., Wei, F., and Xu, K. (2019) · 1908
Earlier work this paper cites.
Revealing the dark secrets of bert
Kovaleva, O., Romanov, A., Rogers, A., and Rumshisky, A. (2019) · 1908
Earlier work this paper cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al. (2019) · 1910
Earlier work this paper cites.
On information and sufficiency
Kullback, S. and Leibler, R. A. (1951) · 1951
Earlier work this paper cites.
Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex
Hubel, D. H. and Wiesel, T. N. (1962) · 1962
Earlier work this paper cites.
Cognitron: A self-organizing multilayered neural network
Fukushima, K. (1975) · 1975
Earlier work this paper cites.
Object recognition with gradient-based learning
LeCun, Y., Haffner, P., Bottou, L., and Bengio, Y. (1999) · 1999
Earlier work this paper cites.
Self-attention attribution: Interpreting information interactions inside transformer
Hao, Y., Dong, L., Wei, F., and Xu, K. (2020) · 2004
Earlier work this paper cites.
Measuring statistical dependence with hilbert-schmidt norms
Gretton, A., Bousquet, O., Smola, A., and Schölkopf, B. (2005) · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Nguyen, T., Raghu, M., and Kornblith, S. (2020) · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012) · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Geiger, A., Lenz, P., Stiller, C., and Urtasun, R. (2013) · 2013
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., and LeCun, Y. (2013) · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., and Darrell, T. (2014) · 2014
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
Dosovitskiy, A., Springenberg, J. T., Riedmiller, M., and Brox, T. (2014) · 2014
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
Eigen, D., Puhrsch, C., and Fergus, R. (2014) · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014) · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K. and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014) · 2014
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T. (2015) · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J. (2015) · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015) · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q. (2016) · 2016
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020) · 2020
Later among the works it cites.
The devil is in the details: Delving into unbiased data processing for human pose estimation
Huang, J., Zhu, Z., Guo, F., and Huang, G. (2020) · 2020
Later among the works it cites.
What is being transferred in transfer learning?
Neyshabur, B., Sedghi, H., and Zhang, C. (2020) · 2020
Later among the works it cites.
Random erasing data augmentation
Zhong, Z., Zheng, L., Kang, G., Li, S., and Yang, Y. (2020) · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F. (2021) · 2021
Later among the works it cites.
Masked-attention mask transformer for universal image segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016) · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A. (2017) · 2017
Cited alongside, same era.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R. (2017) · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. (2017) · 2017
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. (2018) · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., and Girdhar, R. (2021) · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Later among the works it cites.
Bottom-up human pose estimation via disentangled keypoint regression
Geng, Z., Sun, K., Xiao, B., Zhang, Z., and Wang, J. (2021) · 2021
Later among the works it cites.
Simple copy-paste is a strong data augmentation method for instance segmentation
Ghiasi, G., Cui, Y., Srinivas, A., Qian, R., Lin, T.-Y., Cubuk, E. D., Le, Q. V., and Zoph, B. (2021) · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. (2021) · 2021
Later among the works it cites.
Got-10k: A large high-diversity benchmark for generic object tracking in the wild
Huang, L., Zhao, X., and Huang, K. (2021) · 2021
Later among the works it cites.
Localvit: Bringing locality to vision transformers
Li, Y., Zhang, K., Cao, J., Timofte, R., and Van Gool, L. (2021) · 2021
Later among the works it cites.
Swintrack: A simple and strong baseline for transformer tracking
Lin, L., Fan, H., Xu, Y., and Ling, H. (2021) · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A. (2021) · 2021
Later among the works it cites.
Vision transformers for dense prediction
Ranftl, R., Bochkovskiy, A., and Koltun, V. (2021) · 2021
Later among the works it cites.
Concept generalization in visual representation learning
Sariyildiz, M. B., Kalantidis, Y., Larlus, D., and Alahari, K. (2021) · 2021
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. (2021) · 2021
Later among the works it cites.
Hrformer: High-resolution vision transformer for dense predict
Yuan, Y., Fu, R., Huang, L., Lin, W., Zhang, C., Chen, X., and Wang, J. (2021) · 2021
Later among the works it cites.
Bigdetection: A large-scale benchmark for improved object detector pre-training
Cai, L., Zhang, Z., Zhu, Y., Zhang, L., Mu, L., and Xue, X. (2022) · 2022
Closest in time.
Mixformer: End-to-end tracking with iterative mixed attention
Cui, Y., Jiang, C., Wang, L., and Wu, G. (2022) · 2022
Closest in time.
Scaling up your kernels to 31x31: Revisiting large kernel design in cnns
Ding, X., Zhang, X., Zhou, Y., Han, J., Ding, G., and Sun, J. (2022) · 2022
Closest in time.
Global-local path networks for monocular depth estimation with vertical cutdepth
Kim, D., Ga, W., Ahn, P., Joo, D., Chun, S., and Kim, J. (2022) · 2022
Closest in time.
Binsformer: Revisiting adaptive bins for monocular depth estimation
Li, Z., Wang, X., Liu, X., and Jiang, J. (2022) · 2022
Closest in time.
How do vision transformers work?
Park, N. and Kim, S. (2022) · 2022
Closest in time.