Mfas: Multimodal fusion architecture search
Pérez-Rúa, J.-M., Vielzeuf, V., Pateux, S., Baccouche, M., and Jurie, F · 2019
Later among the works it cites.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Cao, J., Gan, Z., Cheng, Y., Yu, L., Chen, Y.-C., and Liu, J · 2020
Later among the works it cites.
Removing bias in multi-modal classifiers: Regularization by maximizing functional entropies
Gat, I., Schwartz, I., Schwing, A., and Hazan, T · 2020
Later among the works it cites.
Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!
Hessel, J. and Lee, L · 2020
Later among the works it cites.
Mmtm: multimodal transfer module for cnn fusion
Joze, H. R. V., Shaban, A., Iuzzolino, M. L., and Koishida, K · 2020
Later among the works it cites.
A closer look at the robustness of vision-and-language pre-trained models
Original
Li, L., Gan, Z., and Liu, J · 2020
Later among the works it cites.
On modality bias in the tvqa dataset
Winterbottom, T., Xiao, S., McLean, A., and Moubayed, N. A · 2020
Later among the works it cites.
Improving the ability of deep networks to use information from multiple views in breast cancer screening
Wu, N., Jastrzębski, S., Park, J., Moy, L., Cho, K., and Geras, K. J · 2020
Later among the works it cites.
Vivit: A video vision transformer
Original
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Later among the works it cites.
High-performance large-scale image recognition without normalization
Brock, A., De, S., Smith, S. L., and Simonyan, K · 2021
Later among the works it cites.
Perceptual score: What data modalities does your model perceive?
Gat, I., Schwartz, I., and Schwing, A · 2021
Later among the works it cites.
From superficial to deep: Language bias driven curriculum learning for visual question answering
Lao, M., Guo, Y., Liu, Y., Chen, W., Pu, N., and Lew, M. S · 2021
Later among the works it cites.
Learning to balance the learning rates between various modalities via adaptive tracking factor
Sun, Y., Mai, S., and Hu, H · 2021
Later among the works it cites.
Hms: Hierarchical modality selectionfor efficient video recognition
Original
Weng, Z., Wu, Z., Li, H., and Jiang, Y.-G · 2021
Later among the works it cites.