Fetching the paper…
Reading the bibliography…
Vision Transformers (ViTs) are becoming more popular and dominating technique for various vision tasks, compare to Convolutional Neural Networks (CNNs).
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” ICML , 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020
2020
Earlier work this paper cites.
H. Du, X. Yu, and L. Zheng, “Vtnet: Visual transformer network for object goal navigation,” in In ICLR , 2020
2020
Earlier work this paper cites.
K. Islam, “Person search: New paradigm of person re-identification: A survey and outlook of recent works,” Image and Vision Computing , vol. 101, p. 103970, 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” ICLR , 2021
2021
Earlier work this paper cites.
Y. Li, K. Zhang, J. Cao, R. Timofte, and L. Van Gool, “Localvit: Bringing locality to vision transformers,” arXiv , 2021
2021
Earlier work this paper cites.
D. Zhou, B. Kang, X. Jin, L. Yang, X. Lian, Q. Hou, and J. Feng, “Deepvit: Towards deeper vision transformer,” arXiv , 2021
2021
Earlier work this paper cites.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” ICCV , 2021
2021
Earlier work this paper cites.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” ICCV , 2021
2021
Earlier work this paper cites.
C.-F. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” ICCV , 2021
2021
Earlier work this paper cites.
G. Yang, H. Tang, M. Ding, N. Sebe, and E. Ricci, “Transformer-based attention networks for continuous pixel-wise prediction,” ICCV , 2021
2021
Earlier work this paper cites.
Y.-F. Wu, J. Yoon, and S. Ahn, “Generative video transformer: Can objects be the words?” in ICML , 2021
2021
Earlier work this paper cites.
S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “Transreid: Transformer-based object re-identification,” arXiv , 2021
2021
Earlier work this paper cites.
Z. Wang, X. Cun, J. Bao, and J. Liu, “Uformer: A general u-shaped transformer for image restoration,” arXiv , 2021
2021
Earlier work this paper cites.
X. Li, Y. Hou, P. Wang, Z. Gao, M. Xu, and W. Li, “Trear: Transformer-based rgb-d egocentric action recognition,” IEEE Transactions on Cognitive and Developmental Systems , 2021
2021
Earlier work this paper cites.
X. Pan, Z. Xia, S. Song, L. E. Li, and G. Huang, “3d object detection with pointformer,” in CVPR , 2021
2021
Earlier work this paper cites.
Y. Gao, M. Zhou, and D. Metaxas, “Utnet: A hybrid transformer architecture for medical image segmentation,” MICCAI , 2021
2021
Earlier work this paper cites.
X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu, “Transformer tracking,” in CVPR , 2021
2021
Earlier work this paper cites.
L. Yuan, Q. Hou, Z. Jiang, J. Feng, and S. Yan, “Volo: Vision outlooker for visual recognition,” arXiv , 2021
2021
Earlier work this paper cites.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr et al. , “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in CVPR , 2021
2021
Earlier work this paper cites.
D. A. Hudson and C. L. Zitnick, “Generative adversarial transformers,” ICML , 2021
2021
Cited alongside, same era.
B. Zhang, S. Gu, B. Zhang, J. Bao, D. Chen, F. Wen, Y. Wang, and B. Guo, “Styleswin: Transformer-based gan for high-resolution image generation,” arXiv , 2021
2021
Cited alongside, same era.
C. Park, Y. Jeong, M. Cho, and J. Park, “Fast point transformer,” arXiv , 2021
2021
Cited alongside, same era.
L. Zou, Z. Huang, N. Gu, and G. Wang, “6d-vit: Category-level 6d object pose estimation via transformer-based instance representation learning,” arXiv , 2021
2021
Cited alongside, same era.
R. Strudel, R. Garcia, I. Laptev, and C. Schmid, “Segmenter: Transformer for semantic segmentation,” ICCV , 2021
2021
Cited alongside, same era.
K. Islam, S. Lee, D. Han, and H. Moon, “Face recognition using shallow age-invariant data,” in 2021 36th International Conference on Image and Vision Computing New Zealand (IVCNZ) . IEEE, 2021, pp. 1–6
2021
Later among the works it cites.
G. Zhang, P. Zhang, J. Qi, and H. Lu, “Hat: Hierarchical aggregation transformers for person re-identification,” arXiv , 2021
2021
Later among the works it cites.
Y. Li, J. He, T. Zhang, X. Liu, Y. Zhang, and F. Wu, “Diverse part discovery: Occluded person re-identification with part-aware transformer,” in CVPR , 2021
2021
Later among the works it cites.
M. Jia, X. Cheng, S. Lu, and J. Zhang, “Learning disentangled representation implicitly via transformer for occluded person re-identification,” arXiv , 2021
2021
Later among the works it cites.
K. Wu, H. Peng, M. Chen, J. Fu, and H. Chao, “Rethinking and improving relative position encoding for vision transformer,” in ICCV , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” arXiv , 2021
2021
Cited alongside, same era.
S. Wu, T. Wu, F. Lin, S. Tian, and G. Guo, “Fully transformer networks for semantic image segmentation,” arXiv , 2021
2021
Cited alongside, same era.
J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, A. L. Yuille, and Y. Zhou, “Transunet: Transformers make strong encoders for medical image segmentation,” arXiv , 2021
2021
Cited alongside, same era.
J. M. J. Valanarasu, P. Oza, I. Hacihaliloglu, and V. M. Patel, “Medical transformer: Gated axial-attention for medical image segmentation,” arXiv , 2021
2021
Cited alongside, same era.
Z. Zhang, B. Sun, and W. Zhang, “Pyramid medical transformer for medical image segmentation,” arXiv , 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” ICCV , 2021
2021
Cited alongside, same era.
H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin-unet: Unet-like pure transformer for medical image segmentation,” arXiv , 2021
2021
Cited alongside, same era.
2021
Later among the works it cites.
D. Coccomini, N. Messina, C. Gennaro, and F. Falchi, “Combining efficientnet and vision transformers for video deepfake detection,” arXiv , 2022
2022
Closest in time.
B. Li, Y. Zhao, Z. Shi, and L. Sheng, “Danceformer: Music conditioned 3d dance generation with parametric motion transformer,” AAAI , 2022
2022
Closest in time.
Z. Sun, Y. Chen, and S. Xiong, “Ssat: A symmetric semantic-aware transformer network for makeup transfer and removal,” AAAI , 2022
2022
Closest in time.
B. Li, C. Zheng, S. Giancola, and B. Ghanem, “Sctn: Sparse convolution-transformer network for scene flow estimation,” AAAI , 2022
2022
Closest in time.
Z. Fan, Z. Song, H. Liu, Z. Lu, J. He, and X. Du, “Svt-net: Super light-weight sparse voxel transformer for large scale place recognition,” AAAI , 2022
2022
Closest in time.
Y. Bai, X. Yang, X. Liu, J. Jiang, Y. Wang, X. Ji, and W. Gao, “Towards end-to-end image compression and analysis with transformers,” AAAI , 2022
2022
Closest in time.
J. He, J.-N. Chen, S. Liu, A. Kortylewski, C. Yang, Y. Bai, C. Wang, and A. Yuille, “Transfg: A transformer architecture for fine-grained recognition,” AAAI , 2022
2022
Closest in time.
Y. Gong, C.-I. J. Lai, Y.-A. Chung, and J. Glass, “Ssast: Self-supervised audio spectrogram transformer,” AAAI , 2022
2022
Closest in time.
Y. Tian, X. Chu, and H. Wang, “Cctrans: Simplifying and improving crowd counting with transformer,” AAAI , 2022
2022
Closest in time.
J. Liang, J. Cao, Y. Fan, K. Zhang, R. Ranjan, Y. Li, R. Timofte, and L. Van Gool, “Vrt: A video restoration transformer,” AAAI , 2022
2022
Closest in time.
S. Woo, J. Park, I. Koo, S. Lee, M. Jeong, and C. Kim, “Explore and match: End-to-end video grounding with transformer,” AAAI , 2022
2022
Closest in time.
S. Yan, X. Xiong, A. Arnab, Z. Lu, M. Zhang, C. Sun, and C. Schmid, “Multiview transformers for video recognition,” arXiv , 2022
2022
Closest in time.
A. Hatamizadeh, D. Yang, H. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” WACV , 2022
2022
Closest in time.
K. Islam, M. Z. Zaheer, and A. Mahmood, “Face pyramid vision transformer-supplementary,” 2022
2022
Closest in time.