Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao, “Simvlm: Simple visual language model pretraining with weak supervision,”
Original
2021
Later among the works it cites.
R. Hu and A. Singh, “Unit: Multimodal multitask learning with a unified transformer,” in
2021
Later among the works it cites.
W. Kim, B. Son, and I. Kim, “Vilt: Vision-and-language transformer without convolution or region supervision,” in
2021
Later among the works it cites.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “Vinvl: Revisiting visual representations in vision-language models,” in
2021
Later among the works it cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in
2021
Later among the works it cites.
H. Li, N. Wang, X. Ding, X. Yang, and X. Gao, “Adaptively learning facial expression representation via cf labels and distillation,”
2021
Later among the works it cites.
W. Zhang, X. Ji, K. Chen, Y. Ding, and C. Fan, “Learning a facial expression embedding disentangled from identity,” in
2021
Later among the works it cites.
Y. Li, Y. Lu, M. Gong, L. Liu, and L. Zhao, “Dual-channel feature disentanglement for identity-invariant facial expression recognition,”
2022
Later among the works it cites.
Q. Zhou, X. Wu, S. Zhang, B. Kang, Z. Ge, and L. J. Latecki, “Contextual ensemble network for semantic segmentation,”
2022
Later among the works it cites.
J. Gu, H. Kwon, D. Wang, W. Ye, M. Li, Y.-H. Chen, L. Lai, V. Chandra, and D. Z. Pan, “Multi-scale high-resolution vision transformer for semantic segmentation,” in
2022
Later among the works it cites.
L. Sun, G. Zhao, Y. Zheng, and Z. Wu, “Spectral–spatial feature tokenization transformer for hyperspectral image classification,”
2022
Later among the works it cites.
G. Han, J. Ma, S. Huang, L. Chen, and S.-F. Chang, “Few-shot object detection with fully cross-transformer,” in
2022
Later among the works it cites.
Z. Tu, H. Talebi, H. Zhang, F. Yang, P. Milanfar, A. Bovik, and Y. Li, “Maxim: Multi-axis mlp for image processing,” in
2022
Later among the works it cites.
J. Shao, Z. Wu, Y. Luo, S. Huang, X. Pu, and Y. Ren, “Self-paced label distribution learning for in-the-wild facial expression recognition,” in
2022
Later among the works it cites.
H. Li, N. Wang, X. Yang, X. Wang, and X. Gao, “Towards semi-supervised deep facial expression recognition with an adaptive confidence margin,” in
2022
Later among the works it cites.
L. Wang, G. Jia, N. Jiang, H. Wu, and J. Yang, “Ease: Robust facial expression recognition via emotion ambiguity-sensitive cooperative networks,” in
2022
Later among the works it cites.
C. Wang, M. Chai, M. He, D. Chen, and J. Liao, “Clip-nerf: Text-and-image driven manipulation of neural radiance fields,” in
2022
Later among the works it cites.
J. Guo, K. Han, H. Wu, Y. Tang, X. Chen, Y. Wang, and C. Xu, “Cmt: Convolutional neural networks meet vision transformers,” in
2022
Later among the works it cites.
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela, “Flava: A foundational language and vision alignment model,” in
2022
Later among the works it cites.
J. Jiang and W. Deng, “Disentangling identity and pose for facial expression recognition,”
2022
Later among the works it cites.
H. Li, N. Wang, X. Yang, and X. Gao, “Crs-cont: A well-trained general encoder for facial expression analysis,”
2022
Later among the works it cites.