Fetching the paper…
Reading the bibliography…
In the past few years, the emergence of pre-training models has brought uni-modal fields such as computer vision (CV) and natural language processing (NLP) to a new era.
Taylor, W.L.: “Cloze procedure”: A new tool for measuring readability. Journalism quarterly 30
1953
Earlier work this paper cites.
Maass, W.: Networks of spiking neurons: the third generation of neural network models. Neural networks 10
1997
Earlier work this paper cites.
Mori, S., Nishida, H., et al
1999
Earlier work this paper cites.
2004
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255 (2009). Ieee
2009
Earlier work this paper cites.
Ordonez, V., Kulkarni, G., et al.: Im2text: Describing images using 1 million captioned photographs. NeurIPS 24
2011
Earlier work this paper cites.
Young, P., Lai, A., et al
2014
Earlier work this paper cites.
Lin, T.-Y., Maire, M., et al
2014
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28
2015
Earlier work this paper cites.
Antol, S., Agrawal, A., et al
2015
Earlier work this paper cites.
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3156–3164 (2015)
2015
Earlier work this paper cites.
Xu, K., Ba, J., et al
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778 (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Specia, L., Frank, S., et al
2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., et al.: Attention is all you need. NeurIPS 30
2017
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299–6308 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Krishna, R., Zhu, Y., et al
2017
Earlier work this paper cites.
Goyal, Y., Khot, T., et al
2017
Earlier work this paper cites.
Chang, A., Dai, A., et al
2017
Earlier work this paper cites.
Wu, Q., Teney, D., Wang, P., Shen, C., Dick, A., Van Den Hengel, A.: Visual question answering: A survey of methods and datasets. Computer Vision and Image Understanding 163
2017
Earlier work this paper cites.
Kafle, K., Kanan, C.: Visual question answering: Datasets, algorithms, and future challenges. Computer Vision and Image Understanding 163
2017
Earlier work this paper cites.
Kafle, K., Kanan, C.: An analysis of visual question answering algorithms. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 1965–1973 (2017)
2017
Earlier work this paper cites.
Suhr, A., Lewis, M., et al
2017
Earlier work this paper cites.
Das, A., Kottur, S., et al
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Anderson, P., He, X., et al
2018
Earlier work this paper cites.
Lei, J., Yu, L., et al
2018
Earlier work this paper cites.
Bai, S., An, S.: A survey on automatic image caption generation. Neurocomputing 311
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Sharma, P., Ding, N., et al
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Zhang, H., Niu, Y., Chang, S.-F.: Grounding referring expressions in images by variational context. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4158–4166 (2018)
2018
Earlier work this paper cites.
Ghosal, D., Akhtar, M.S., Chauhan, D., Poria, S., Ekbal, A., Bhattacharyya, P.: Contextual inter-modal attention for multi-modal sentiment analysis. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 3454–3466 (2018)
2018
Earlier work this paper cites.
Mithun, N.C., Li, J., Metze, F., Roy-Chowdhury, A.K.: Learning joint embedding with multimodal cues for cross-modal video-text retrieval. In: Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval, pp. 19–27 (2018)
2018
Earlier work this paper cites.
Wang, B., Ma, L., Zhang, W., Liu, W.: Reconstruction network for video captioning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7622–7631 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Devlin, J., Chang, M., et al
2019
Earlier work this paper cites.
Schneider, S., Baevski, A., et al
2019
Earlier work this paper cites.
Lu, J., Batra, D., et al
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6202–6211 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Lan, Z., Chen, M., et al
2019
Earlier work this paper cites.
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R.R., Le, Q.V.: Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems 32
2019
Earlier work this paper cites.
Hudson, D.A., Manning, C.D.: Gqa: A new dataset for real-world visual reasoning and compositional question answering. In: CVPR, pp. 6700–6709 (2019)
2019
Earlier work this paper cites.
Miech, A., Zhukov, D., et al
2019
Earlier work this paper cites.
Sun, C., Myers, A., et al
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Zellers, R., Bisk, Y., et al
2019
Earlier work this paper cites.
Yu, W., Zhou, J., Yu, W., Liang, X., Xiao, N.: Heterogeneous graph learning for visual commonsense reasoning. Advances in Neural Information Processing Systems 32
2019
Earlier work this paper cites.
Liu, X., Wang, Z., et al
2019
Cited alongside, same era.
Yang, S., Li, G., Yu, Y.: Cross-modal relationship inference for grounding referring expressions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4145–4154 (2019)
2019
Cited alongside, same era.
Akhtar, M.S., Chauhan, D., Ghosal, D., Poria, S., Ekbal, A., Bhattacharyya, P.: Multi-task Learning for Multi-modal Emotion Recognition and Sentiment Analysis. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 370–379 (2019)
2019
Cited alongside, same era.
Agrawal, H., Desai, K., et al
2019
Cited alongside, same era.
Su, Y., Fan, K., Bach, N., Kuo, C.-C.J., Huang, F.: Unsupervised multi-modal neural machine translation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10482–10491 (2019)
Touvron, H., Cord, M., et al
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Kamath, A., Singh, M., et al
2021
Later among the works it cites.
Zhuge, M., Gao, D., et al
2021
Later among the works it cites.
Changpinyo, S., Sharma, P., et al
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Wang, X., Huang, Q., Celikyilmaz, A., Gao, J., Shen, D., Wang, Y.-F., Wang, W.Y., Zhang, L.: Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6629–6638 (2019)
2019
Cited alongside, same era.
Zhao, Z.-Q., Zheng, P., Xu, S.-t., Wu, X.: Object detection with deep learning: A review. IEEE transactions on neural networks and learning systems 30
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Tan, H., Bansal, M.: LXMERT: Learning cross-modality encoder representations from transformers. In: EMNLP, pp. 5100–5111 (2019)
2019
Cited alongside, same era.
Alberti, C., Ling, J., et al
2019
Cited alongside, same era.
Su, W., Zhu, X., et al
2019
Cited alongside, same era.
Fan, A., Grave, E., et al
2019
Cited alongside, same era.
Jia, C., Yang, Y., et al
2021
Later among the works it cites.
Bain, M., Nagrani, A., et al
2021
Later among the works it cites.
Bitton, Y., Stanovsky, G., Schwartz, R., Elhadad, M.: Automatic Generation of Contrast Sets from Scene Graphs: Probing the Compositional Consistency of GQA. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 94–105 (2021)
2021
Later among the works it cites.
Li, J., Tang, S., Zhu, L., Shi, H., Huang, X., Wu, F., Yang, Y., Zhuang, Y.: Adaptive hierarchical graph reasoning with semantic coherence for video-and-language inference. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1867–1877 (2021)
2021
Later among the works it cites.
Chaudhary, A.: Robust Vision and Language Inference via Semantics Transformed Adversarial Training. PhD thesis, Arizona State University (2021)
2021
Later among the works it cites.
Ye, K., Kovashka, A.: A case study of the shortcut effects in visual commonsense reasoning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 3181–3189 (2021)
2021
Later among the works it cites.
Jiming, L., Peixiang, Z., et al
2021
Later among the works it cites.
Chen, F., Chen, X., Meng, F., Li, P., Zhou, J.: GoG: Relation-aware Graph-over-Graph Network for Visual Dialog. In: Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 230–243 (2021)
2021
Later among the works it cites.
Chen, F., Meng, F., Chen, X., Li, P., Zhou, J.: Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation. In: ACL/IJCNLP (Findings) (2021)
2021
Later among the works it cites.
Strudel, R., Garcia, R., Laptev, I., Schmid, C.: Segmenter: Transformer for semantic segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7262–7272 (2021)
2021
Later among the works it cites.
Fang, Y., Liao, B., Wang, X., Fang, J., Qi, J., Wu, R., Niu, J., Liu, W.: You only look at one sequence: Rethinking transformer in vision through object detection. Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Hong, Y., Wu, Q., et al
2021
Later among the works it cites.
Chiou, M.-J., Zimmermann, R., et al
2021
Later among the works it cites.
Cho, J., Lei, J., et al
2021
Later among the works it cites.
2021
Later among the works it cites.
Huang, Z., Zeng, Z., et al
2021
Later among the works it cites.
Xu, H., Yan, M., et al
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Dong, J., Cong, Y., Sun, G., Fang, Z., Ding, Z.: Where and how to transfer: knowledge aggregation-induced transferability perception for unsupervised domain adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021)
2021
Later among the works it cites.
Zhu, H., Luo, M.-D., Wang, R., Zheng, A.-H., He, R.: Deep audio-visual learning: A survey. International Journal of Automation and Computing 18
2021
Later among the works it cites.
Tao, J.-H., Huang, J., Li, Y., Lian, Z., Niu, M.-Y.: Correction to: Semi-supervised ladder networks for speech emotion recognition. International Journal of Automation and Computing 18
2021
Later among the works it cites.
Akbari, H., Yuan, L., et al.: VATT: Transformers for multimodal self-supervised learning from raw video, audio and text. NeurIPS 34
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Chen, K., Huang, Q., et al
2021
Later among the works it cites.
Tsimpoukelli, M., Menick, J., et al.: Multimodal few-shot learning with frozen language models. NeurIPS 34
2021
Later among the works it cites.
Fang, Z., Wang, J., et al
2021
Later among the works it cites.
Li, Y., Liang, F., et al
2021
Later among the works it cites.
2021
Later among the works it cites.
Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y., et al.: A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)
2022
Closest in time.
Chen, F., Chen, X., Xu, S., Xu, B.: Improving cross-modal understanding in visual dialog via contrastive learning. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7937–7941 (2022). IEEE
2022
Closest in time.
Song, H., Dong, L., Zhang, W., Liu, T., Wei, F.: CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6088–6100 (2022)
2022
Closest in time.
2022
Closest in time.
Gu, J., Stefani, E., et al
2022
Closest in time.
Mo, Y., Wu, Y., Yang, X., Liu, F., Liao, Y.: Review the state-of-the-art technologies of semantic segmentation based on deep learning. Neurocomputing 493
2022
Closest in time.
Rethmeier, N., Augenstein, I.: Long-tail zero and few-shot learning via contrastive pretraining on and for small data. In: Computer Sciences & Mathematics Forum, vol. 3, p. 10 (2022). MDPI
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Chen, W., Han, X., Lin, Y., Zhao, H., Liu, Z., Li, P., Sun, M., Zhou, J.: Fully Hyperbolic Neural Networks. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5672–5686 (2022)
2022
Closest in time.