Fetching the paper…
Reading the bibliography…
Transformer is a promising neural network learner, and has achieved great success in various machine learning tasks.
A. A. Lazarus et al. , Multimodal behavior therapy . Springer, 1976
1976
Earlier work this paper cites.
B. P. Yuhas, M. H. Goldstein, and T. J. Sejnowski, “Integration of acoustic and visual speech signals using neural networks,” IEEE Communications Magazine , 1989
1989
Earlier work this paper cites.
T. Chen and R. R. Rao, “Audio-visual integration in multimodal communication,” Proceedings of the IEEE , 1998
1998
Earlier work this paper cites.
L. Wu, S. L. Oviatt, and P. R. Cohen, “Multimodal integration-a statistical view,” TMM , 1999
1999
Earlier work this paper cites.
T. Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman, The elements of statistical learning: data mining, inference, and prediction . Springer, 2009, vol. 2
2009
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” NeurIPS , 2011
2011
Earlier work this paper cites.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in AISTATS , 2011
2011
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in ICCV , 2015
2015
Earlier work this paper cites.
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy, “What’s cookin’? interpreting cooking videos using text, speech and vision,” arXiv , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” NeurIPS , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv , 2016
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , 2017
2017
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in ICCV , 2017
2017
Earlier work this paper cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra, “Visual dialog,” in CVPR , 2017
2017
Earlier work this paper cites.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning for video understanding,” arXiv , 2017
2017
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” TPAMI , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv , 2018
2018
Earlier work this paper cites.
P. Xu, Y. Huang, T. Yuan, K. Pang, Y.-Z. Song, T. Xiang, T. M. Hospedales, Z. Ma, and J. Guo, “Sketchmate: Deep hashing for million-scale human sketch retrieval,” in CVPR , 2018
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL , 2018
2018
Earlier work this paper cites.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,” TPAMI , 2018
2018
Earlier work this paper cites.
L. Zhou, C. Xu, and J. J. Corso, “Towards automatic learning of procedures from web instructional videos,” in AAAI , 2018
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR , 2018
2018
Earlier work this paper cites.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in CVPR , 2018
2018
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in CVPR , 2018
2018
Earlier work this paper cites.
M. Chen, Y. Li, Z. Zhang, and S. Huang, “Tvt: Two-view transformer network for video captioning,” in ACML , 2018
2018
Earlier work this paper cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in ECCV , 2018
2018
Earlier work this paper cites.
Y. Xian, C. H. Lampert, B. Schiele, and Z. Akata, “Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,” TPAMI , 2018
2018
Earlier work this paper cites.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in ICCV , 2019
2019
Earlier work this paper cites.
D. Shin, Z. Ren, E. B. Sudderth, and C. C. Fowlkes, “3d scene reconstruction with multi-layer depth and epipolar transformers,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Shang, T. Ma, C. Xiao, and J. Sun, “Pre-training of graph augmented transformers for medication recommendation,” arXiv , 2019
2019
Earlier work this paper cites.
W. Guo, J. Wang, and S. Wang, “Deep multimodal representation learning: A survey,” IEEE Access , 2019
2019
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” arXiv , 2019
2019
Earlier work this paper cites.
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” arXiv , 2019
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” NeurIPS , 2019
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” arXiv , 2019
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” arXiv , 2019
2019
Earlier work this paper cites.
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang, “Visualbert: A simple and performant baseline for vision and language,” arXiv , 2019
2019
Earlier work this paper cites.
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai, “Vl-bert: Pre-training of generic visual-linguistic representations,” arXiv , 2019
2019
Earlier work this paper cites.
C. Sun, F. Baradel, K. Murphy, and C. Schmid, “Learning video representations using contrastive bidirectional transformer,” arXiv , 2019
2019
Earlier work this paper cites.
C. Alberti, J. Ling, M. Collins, and D. Reitter, “Fusion of detected objects in text for visual question answering,” arXiv , 2019
2019
Earlier work this paper cites.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” in ICCV , 2019
2019
Earlier work this paper cites.
B. Wang, R. Shin, X. Liu, O. Polozov, and M. Richardson, “Rat-sql: Relation-aware schema encoding and linking for text-to-sql parsers,” arXiv , 2019
2019
Earlier work this paper cites.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in ACL , 2019
2019
Earlier work this paper cites.
L. Huang, W. Wang, J. Chen, and X.-Y. Wei, “Attention on attention for image captioning,” in ICCV , 2019
2019
Earlier work this paper cites.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in AAAI , 2019
2019
Earlier work this paper cites.
R. Child, S. Gray, A. Radford, and I. Sutskever, “Generating long sequences with sparse transformers,” arXiv , 2019
2019
Earlier work this paper cites.
S. Pramanik, P. Agrawal, and A. Hussain, “Omninet: A unified architecture for multi-modal multi-task learning,” arXiv , 2019
2019
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv , 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020
2020
Earlier work this paper cites.
C. Zhang, Z. Yang, X. He, and L. Deng, “Multimodal intelligence: Representation learning, information fusion, and applications,” JSTSP , 2020
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” NeurIPS , 2020
2020
Earlier work this paper cites.
A. Nagrani, C. Sun, D. Ross, R. Sukthankar, C. Schmid, and A. Zisserman, “Speech2action: Cross-modal supervision for action recognition,” in CVPR , 2020
2020
Earlier work this paper cites.
W. Chen, M.-W. Chang, E. Schlinger, W. Wang, and W. W. Cohen, “Open question answering over tables and text,” arXiv , 2020
2020
Earlier work this paper cites.
K. Gupta, J. Lazarow, A. Achille, L. Davis, V. Mahadevan, and A. Shrivastava, “Layouttransformer: Layout generation and completion with self-attention,” arXiv , 2020
2020
Earlier work this paper cites.
P. Xu, Z. Song, Q. Yin, Y.-Z. Song, and L. Wang, “Deep self-supervised representation learning for free-hand sketch,” TCSVT , 2020
2020
Earlier work this paper cites.
P. Xu, K. Liu, T. Xiang, T. M. Hospedales, Z. Ma, J. Guo, and Y.-Z. Song, “Fine-grained instance-level sketch-based video retrieval,” TCSVT , 2020
2020
Earlier work this paper cites.
Y. Xu, Y. Xu, T. Lv, L. Cui, F. Wei, G. Wang, Y. Lu, D. Florencio, C. Zhang, W. Che et al. , “Layoutlmv2: Multi-modal pre-training for visually-rich document understanding,” arXiv , 2020
2020
Earlier work this paper cites.
I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The long-document transformer,” arXiv , 2020
2020
Earlier work this paper cites.
D. Guo, S. Ren, S. Lu, Z. Feng, D. Tang, S. Liu, L. Zhou, N. Duan, A. Svyatkovskiy, S. Fu et al. , “Graphcodebert: Pre-training code representations with data flow,” arXiv , 2020
2020
Earlier work this paper cites.
K. Gavrilyuk, R. Sanford, M. Javan, and C. G. Snoek, “Actor-transformers for group activity recognition,” in CVPR , 2020
2020
Earlier work this paper cites.
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient transformers: A survey,” arXiv , 2020
2020
Earlier work this paper cites.
A. M. Braşoveanu and R. Andonie, “Visualizing transformers for nlp: a brief survey,” in International Conference Information Visualisation (IV) , 2020
2020
Earlier work this paper cites.
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu et al. , “A survey on vision transformer,” arXiv , 2020
2020
Earlier work this paper cites.
D. Feng, C. Haase-Schütz, L. Rosenbaum, H. Hertlein, C. Glaeser, F. Timm, W. Wiesbeck, and K. Dietmayer, “Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges,” TITS , 2020
2020
Earlier work this paper cites.
Y. Li, S. Rao, J. R. A. Solares, A. Hassaine, R. Ramakrishnan, D. Canoy, Y. Zhu, K. Rahimi, and G. Salimi-Khorshidi, “Behrt: transformer for electronic health records,” Scientific reports , 2020
2020
Earlier work this paper cites.
Y. Li, H. Wang, and Y. Luo, “A comparison of pre-trained vision-and-language models for multimodal representation learning across medical images and reports,” in BIBM , 2020
2020
Earlier work this paper cites.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in ICML , 2020
2020
Earlier work this paper cites.
J. Beal, E. Kim, E. Tzeng, D. H. Park, A. Zhai, and D. Kislyuk, “Toward transformer-based object detection,” arXiv , 2020
2020
Earlier work this paper cites.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” in ECCV , 2020
2020
Earlier work this paper cites.
G. Li, N. Duan, Y. Fang, M. Gong, and D. Jiang, “Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,” in AAAI , 2020
2020
Earlier work this paper cites.
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao, “Unified vision-language pre-training for image captioning and vqa,” in AAAI , 2020
2020
Earlier work this paper cites.
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee, “12-in-1: Multi-task vision and language representation learning,” in CVPR , 2020
2020
Earlier work this paper cites.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei et al. , “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in ECCV , 2020
2020
Earlier work this paper cites.
Z. Huang, Z. Zeng, B. Liu, D. Fu, and J. Fu, “Pixel-bert: Aligning image pixels with text by deep multi-modal transformers,” arXiv , 2020
2020
Earlier work this paper cites.
L. Zhu and Y. Yang, “Actbert: Learning global-local video-text representations,” in CVPR , 2020
2020
Earlier work this paper cites.
D. Qi, L. Su, J. Song, E. Cui, T. Bharti, and A. Sacheti, “Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data,” arXiv , 2020
2020
Earlier work this paper cites.
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu, “Hero: Hierarchical encoder for video+ language omni-representation pre-training,” arXiv , 2020
2020
Earlier work this paper cites.
H. Luo, L. Ji, B. Shi, H. Huang, N. Duan, T. Li, J. Li, T. Bharti, and M. Zhou, “Univl: A unified video and language pre-training model for multimodal understanding and generation,” arXiv , 2020
2020
Earlier work this paper cites.
P. Morgado, Y. Li, and N. Vasconcelos, “Learning representations from audio-visual spatial alignment,” arXiv , 2020
2020
Earlier work this paper cites.
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, and D. Testuggine, “The hateful memes challenge: Detecting hate speech in multimodal memes,” arXiv , 2020
2020
Earlier work this paper cites.
V. P. Dwivedi and X. Bresson, “A generalization of transformer networks to graphs,” arXiv , 2020
2020
Earlier work this paper cites.
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu, “On layer normalization in the transformer architecture,” in ICML , 2020
2020
Earlier work this paper cites.
J. Lin, A. Yang, Y. Zhang, J. Liu, J. Zhou, and H. Yang, “Interbert: Vision-and-language interaction for multi-modal pretraining,” arXiv , 2020
2020
Earlier work this paper cites.
V. Murahari, D. Batra, D. Parikh, and A. Das, “Large-scale pretraining for visual dialog: A simple state-of-the-art baseline,” in ECCV , 2020
2020
Earlier work this paper cites.
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman, “End-to-end learning of visual representations from uncurated instructional videos,” in CVPR , 2020
2020
Earlier work this paper cites.
J. Chen, M. Ma, R. Zheng, and L. Huang, “Mam: Masked acoustic modeling for end-to-end speech-to-text translation,” arXiv , 2020
2020
Earlier work this paper cites.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in CVPR , 2020
2020
Earlier work this paper cites.
Z. Gan, Y.-C. Chen, L. Li, C. Zhu, Y. Cheng, and J. Liu, “Large-scale adversarial training for vision-and-language representation learning,” arXiv , 2020
2020
Earlier work this paper cites.
V. Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal transformer for video retrieval,” in ECCV , 2020
2020
Earlier work this paper cites.
X. Lin, C. Ding, J. Zeng, and D. Tao, “Gps-net: Graph property sensing network for scene graph generation,” in CVPR , 2020
2020
Earlier work this paper cites.
Y. Pan, T. Yao, Y. Li, and T. Mei, “X-linear attention networks for image captioning,” in CVPR , 2020
2020
Earlier work this paper cites.
R. Huang, H. Hu, W. Wu, K. Sawada, M. Zhang, and D. Jiang, “Dance revolution: Long-term dance generation with music via curriculum learning,” arXiv , 2020
2020
Earlier work this paper cites.
M. Chen, X. Tan, Y. Ren, J. Xu, H. Sun, S. Zhao, T. Qin, and T.-Y. Liu, “Multispeech: Multi-speaker text to speech with transformer,” arXiv , 2020
2020
Earlier work this paper cites.
S. Ging, M. Zolfaghari, H. Pirsiavash, and T. Brox, “Coot: Cooperative hierarchical transformer for video-text representation learning,” arXiv , 2020
2020
Earlier work this paper cites.
M. Patrick, P.-Y. Huang, Y. Asano, F. Metze, A. Hauptmann, J. Henriques, and A. Vedaldi, “Support-set bottlenecks for video-text representation learning,” arXiv , 2020
2020
Earlier work this paper cites.
A. Sadhu, K. Chen, and R. Nevatia, “Video object grounding using semantic roles in language description,” in CVPR , 2020
2020
Earlier work this paper cites.
J. Cho, J. Lu, D. Schwenk, H. Hajishirzi, and A. Kembhavi, “X-lxmert: Paint, caption and answer questions with multi-modal transformers,” in EMNLP , 2020
2020
Earlier work this paper cites.
W. Rahman, M. K. Hasan, S. Lee, A. Zadeh, C. Mao, L.-P. Morency, and E. Hoque, “Integrating multimodal information in large pretrained transformers,” in ACL , 2020
2020
Earlier work this paper cites.
S. Lee, Y. Yu, G. Kim, T. Breuel, J. Kautz, and Y. Song, “Parameter efficient multimodal transformers for video representation learning,” arXiv , 2020
2020
Cited alongside, same era.
W. Wang and Z. Tu, “Rethinking the value of transformer components,” in COLING , 2020
2020
Cited alongside, same era.
A. R. Akula, S. Gella, Y. Al-Onaizan, S.-C. Zhu, and S. Reddy, “Words aren’t enough, their order matters: On the robustness of grounding visual referring expressions,” arXiv , 2020
2020
Cited alongside, same era.
L. Li, Z. Gan, and J. Liu, “A closer look at the robustness of vision-and-language pre-trained models,” arXiv , 2020
2020
Cited alongside, same era.
J. Cao, Z. Gan, Y. Cheng, L. Yu, Y.-C. Chen, and J. Liu, “Behind the scene: Revealing the secrets of pre-trained vision-and-language models,” in ECCV , 2020
2020
Cited alongside, same era.
H. Zhang, W. Yin, Y. Fang, L. Li, B. Duan, Z. Wu, Y. Sun, H. Tian, H. Wu, and H. Wang, “Ernie-vilg: Unified generative pre-training for bidirectional vision-language generation,” arXiv , 2021
2021
Later among the works it cites.
M. Zhuge, D. Gao, D.-P. Fan, L. Jin, B. Chen, H. Zhou, M. Qiu, and L. Shao, “Kaleido-bert: Vision-language pre-training on fashion domain,” in CVPR , 2021
2021
Later among the works it cites.
A. Prakash, K. Chitta, and A. Geiger, “Multi-modal fusion transformer for end-to-end autonomous driving,” in CVPR , 2021
2021
Later among the works it cites.
X. Favory, K. Drossos, T. Virtanen, and X. Serra, “Learning contextual tag embeddings for cross-modal alignment of audio and tags,” in ICASSP , 2021
2021
Later among the works it cites.
A. Botach, E. Zheltonozhskii, and C. Baskin, “End-to-end referring video object segmentation with multimodal transformers,” arXiv , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira, “Perceiver: General perception with iterative attention,” in ICML , 2021
2021
Cited alongside, same era.
J. Chen, X. Tan, Y. Leng, J. Xu, G. Wen, T. Qin, and T.-Y. Liu, “Speech-t: Transducer for text to speech and beyond,” NeurIPS , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” arXiv , 2021
2021
Cited alongside, same era.
F. Qingyun, H. Dapeng, and W. Zhaokui, “Cross-modality fusion transformer for multispectral object detection,” arXiv , 2021
2021
Cited alongside, same era.
Y. Guo, L. Gao, X. Wang, Y. Hu, X. Xu, X. Lu, H. T. Shen, and J. Song, “From general to specific: Informative scene graph generation via balance adjustment,” in ICCV , 2021
2021
Cited alongside, same era.
C.-F. Yang, W.-C. Fan, F.-E. Yang, and Y.-C. F. Wang, “Layouttransformer: Scene layout generation with conceptual and spatial diversity,” in CVPR , 2021
2021
Cited alongside, same era.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in CVPR , 2021
2021
Cited alongside, same era.
2021
Later among the works it cites.
W. Hong, K. Ji, J. Liu, J. Wang, J. Chen, and W. Chu, “Gilbert: Generative vision-language pre-training for image-text retrieval,” in SIGIR , 2021
2021
Later among the works it cites.
X. Xu and C. C. Loy, “3d human texture estimation from a single image with transformers,” in ICCV , 2021
2021
Later among the works it cites.
W. Wang, R. Wang, and X. Chen, “Topic scene graph generation by attention distillation from caption,” in ICCV , 2021
2021
Later among the works it cites.
Y. Lu, H. Rai, J. Chang, B. Knyazev, G. Yu, S. Shekhar, G. W. Taylor, and M. Volkovs, “Context-aware scene graph generation with seq2seq transformers,” in ICCV , 2021
2021
Later among the works it cites.
P. Ke, H. Ji, Y. Ran, X. Cui, L. Wang, L. Song, X. Zhu, and M. Huang, “Jointgt: Graph-text joint representation learning for text generation from knowledge graphs,” arXiv , 2021
2021
Later among the works it cites.
Y. Teng, L. Wang, Z. Li, and G. Wu, “Target adaptive context aggregation for video scene graph generation,” in ICCV , 2021
2021
Later among the works it cites.
K. Lin, L. Li, C.-C. Lin, F. Ahmed, Z. Gan, Z. Liu, Y. Lu, and L. Wang, “Swinbert: End-to-end transformers with sparse attention for video captioning,” arXiv , 2021
2021
Later among the works it cites.
C. Deng, S. Chen, D. Chen, Y. He, and Q. Wu, “Sketch, ground, and refine: Top-down dense video captioning,” in CVPR , 2021
2021
Later among the works it cites.
T. Wang, R. Zhang, Z. Lu, F. Zheng, R. Cheng, and P. Luo, “End-to-end dense video captioning with parallel decoding,” in ICCV , 2021
2021
Later among the works it cites.
X. Yang, H. Zhang, G. Qi, and J. Cai, “Causal attention for vision-language tasks,” in CVPR , 2021
2021
Later among the works it cites.
Y. Luo, J. Ji, X. Sun, L. Cao, Y. Wu, F. Huang, C.-W. Lin, and R. Ji, “Dual-level collaborative transformer for image captioning,” arXiv , 2021
2021
Later among the works it cites.
G. Xu, S. Niu, M. Tan, Y. Luo, Q. Du, and Q. Wu, “Towards accurate text-based image captioning with content diversity exploration,” in CVPR , 2021
2021
Later among the works it cites.
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang et al. , “Cogview: Mastering text-to-image generation via transformers,” arXiv , 2021
2021
Later among the works it cites.
A. Sanghi, H. Chu, J. G. Lambourne, Y. Wang, C.-Y. Cheng, and M. Fumero, “Clip-forge: Towards zero-shot text-to-shape generation,” arXiv , 2021
2021
Later among the works it cites.
Y. Zhong, J. Shi, J. Yang, C. Xu, and Y. Li, “Learning to generate scene graph from natural language supervision,” in ICCV , 2021
2021
Later among the works it cites.
S. Geng, P. Gao, M. Chatterjee, C. Hori, J. Le Roux, Y. Zhang, H. Li, and A. Cherian, “Dynamic graph representation learning for video dialog via multi-modal shuffled transformers,” in AAAI , 2021
2021
Later among the works it cites.
T. Jaunet, C. Kervadec, R. Vuillemot, G. Antipov, M. Baccouche, and C. Wolf, “Visqa: X-raying vision and language reasoning in transformers,” TVCG , 2021
2021
Later among the works it cites.
X. Lin, G. Bertasius, J. Wang, S.-F. Chang, D. Parikh, and L. Torresani, “Vx2text: End-to-end learning of video-based text generation from multimodal inputs,” in CVPR , 2021
2021
Later among the works it cites.
J. Lin, R. Men, A. Yang, C. Zhou, Y. Zhang, P. Wang, J. Zhou, J. Tang, and H. Yang, “M6: Multi-modality-to-multi-modality multitask mega-transformer for unified pretraining,” in KDD , 2021
2021
Later among the works it cites.
H. Xue, Y. Huang, B. Liu, H. Peng, J. Fu, H. Li, and J. Luo, “Probing inter-modality: Visual parsing with self-attention for vision-and-language pre-training,” NeurIPS , 2021
2021
Later among the works it cites.
T.-D. Truong, C. N. Duong, H. A. Pham, B. Raj, N. Le, K. Luu et al. , “The right to talk: An audio-visual transformer approach,” in ICCV , 2021
2021
Later among the works it cites.
Y. Zhang, M. Choi, K. Han, and Z. Liu, “Explainable semantic space by grounding language to vision with cross-modal contrastive learning,” NeurIPS , 2021
2021
Later among the works it cites.
Y.-W. Chen, Y.-H. Tsai, and M.-H. Yang, “End-to-end multi-modal video temporal grounding,” NeurIPS , 2021
2021
Later among the works it cites.
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer, “Videoclip: Contrastive pre-training for zero-shot video-text understanding,” arXiv , 2021
2021
Later among the works it cites.
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less is more: Clipbert for video-and-language learning via sparse sampling,” in CVPR , 2021
2021
Later among the works it cites.
J. Yang, Y. Bisk, and J. Gao, “Taco: Token-aware cascade contrastive learning for video-text alignment,” in ICCV , 2021
2021
Later among the works it cites.
D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi, “Align and prompt: Video-and-language pre-training with entity prompts,” arXiv , 2021
2021
Later among the works it cites.
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “Clip4clip: An empirical study of clip for end to end video clip retrieval,” arXiv , 2021
2021
Later among the works it cites.
H. Fang, P. Xiong, L. Xu, and Y. Chen, “Clip2video: Mastering video-text retrieval via image clip,” arXiv , 2021
2021
Later among the works it cites.
M. Narasimhan, A. Rohrbach, and T. Darrell, “Clip-it! language-guided video summarization,” NeurIPS , 2021
2021
Later among the works it cites.
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui, “Open-vocabulary object detection via vision and language knowledge distillation,” arXiv , 2021
2021
Later among the works it cites.
H. Xu, M. Yan, C. Li, B. Bi, S. Huang, W. Xiao, and F. Huang, “E2e-vlp: End-to-end vision-language pre-training enhanced by visual learning,” arXiv , 2021
2021
Later among the works it cites.
C. Kervadec, C. Wolf, G. Antipov, M. Baccouche, and M. Nadri, “Supervising the transfer of reasoning patterns in vqa,” arXiv , 2021
2021
Later among the works it cites.
C. Kervadec, T. Jaunet, G. Antipov, M. Baccouche, R. Vuillemot, and C. Wolf, “How transferable are reasoning patterns in vqa?” in CVPR , 2021
2021
Later among the works it cites.
D. Agarwal, T. Agrawal, L. M. Ferrari, and F. Bremond, “From multimodal to unimodal attention in transformers using knowledge distillation,” in AVSS , 2021
2021
Later among the works it cites.
Q. Li, B. Gong, Y. Cui, D. Kondratyuk, X. Du, M.-H. Yang, and M. Brown, “Towards a unified foundation model: Jointly pre-training transformers on unpaired images and text,” arXiv , 2021
2021
Later among the works it cites.
M. Ni, H. Huang, L. Su, E. Cui, T. Bharti, L. Wang, D. Zhang, and N. Duan, “M3p: Learning universal representations via multitask multilingual multimodal pre-training,” in CVPR , 2021
2021
Later among the works it cites.
A. Miech, J.-B. Alayrac, I. Laptev, J. Sivic, and A. Zisserman, “Thinking fast and slow: Efficient text-to-visual retrieval with transformers,” in CVPR , 2021
2021
Later among the works it cites.
K. Wen, J. Xia, Y. Huang, L. Li, J. Xu, and J. Shao, “Cookie: Contrastive cross-modal knowledge sharing pre-training for vision-language representation,” in ICCV , 2021
2021
Later among the works it cites.
T. Liu, F. Feng, and X. Wang, “Multi-stage pre-training over simplified multimodal pre-training models,” arXiv , 2021
2021
Later among the works it cites.
Y. Li, F. Liang, L. Zhao, Y. Cui, W. Ouyang, J. Shao, F. Yu, and J. Yan, “Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm,” arXiv , 2021
2021
Later among the works it cites.
Z. Gan, Y.-C. Chen, L. Li, T. Chen, Y. Cheng, S. Wang, and J. Liu, “Playing lottery tickets with vision and language,” arXiv , 2021
2021
Later among the works it cites.
C. He, S. Li, M. Soltanolkotabi, and S. Avestimehr, “Pipetransformer: Automated elastic pipelining for distributed training of large-scale models,” in ICML , 2021
2021
Later among the works it cites.
C. Zhu, W. Ping, C. Xiao, M. Shoeybi, T. Goldstein, A. Anandkumar, and B. Catanzaro, “Long-short transformer: Efficient transformers for language and vision,” in NeurIPS , 2021
2021
Later among the works it cites.
A. Akula, V. Jampani, S. Changpinyo, and S.-C. Zhu, “Robust visual reasoning via language guided neural module networks,” NeurIPS , 2021
2021
Later among the works it cites.
M. Zhang, T. Maidment, A. Diab, A. Kovashka, and R. Hwa, “Domain-robust vqa with diverse datasets and methods but no target labels,” in CVPR , 2021
2021
Later among the works it cites.
Y. Kant, A. Moudgil, D. Batra, D. Parikh, and H. Agrawal, “Contrast and classify: Training robust vqa models,” in ICCV , 2021
2021
Later among the works it cites.
L. A. Hendricks, J. Mellor, R. Schneider, J.-B. Alayrac, and A. Nematzadeh, “Decoupling the role of data, attention, and losses in multimodal transformers,” TACL , 2021
2021
Later among the works it cites.
L. A. Hendricks and A. Nematzadeh, “Probing image-language transformers for verb understanding,” arXiv , 2021
2021
Later among the works it cites.
S. Frank, E. Bugliarello, and D. Elliott, “Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers,” arXiv , 2021
2021
Later among the works it cites.
H. Chefer, S. Gur, and L. Wolf, “Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers,” arXiv , 2021
2021
Later among the works it cites.
L. Parcalabescu, M. Cafagna, L. Muradjan, A. Frank, I. Calixto, and A. Gatt, “Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena,” arXiv , 2021
2021
Later among the works it cites.
N. Mu, A. Kirillov, D. Wagner, and S. Xie, “Slip: Self-supervision meets language-image pre-training,” arXiv , 2021
2021
Later among the works it cites.
H. Xu, G. Ghosh, P.-Y. Huang, P. Arora, M. Aminzadeh, C. Feichtenhofer, F. Metze, and L. Zettlemoyer, “Vlm: Task-agnostic video-language model pre-training for video understanding,” arXiv , 2021
2021
Later among the works it cites.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “Vinvl: Revisiting visual representations in vision-language models,” in CVPR , 2021
2021
Later among the works it cites.
M. Li, R. Xu, S. Wang, L. Zhou, X. Lin, C. Zhu, M. Zeng, H. Ji, and S.-F. Chang, “Clip-event: Connecting text and images with event structures,” arXiv , 2022
2022
Closest in time.
A. Rahate, R. Walambe, S. Ramanna, and K. Kotecha, “Multimodal co-learning: Challenges, applications with datasets, recent advances and future directions,” Information Fusion , 2022
2022
Closest in time.
K. K. Parida, S. Srivastava, and G. Sharma, “Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal attention,” in WACV , 2022
2022
Closest in time.
R. Li, S. Zhang, and X. He, “Sgtr: End-to-end scene graph generation with transformer,” in CVPR , 2022
2022
Closest in time.
Y. Song, R. C.-W. Wong, X. Zhao, and D. Jiang, “Speech-to-sql: Towards speech-driven sql query generation from natural language question,” arXiv , 2022
2022
Closest in time.
X. Zhu, Z. Li, X. Wang, X. Jiang, P. Sun, X. Wang, Y. Xiao, and N. J. Yuan, “Multi-modal knowledge graph construction and application: A survey,” arXiv , 2022
2022
Closest in time.
Y. Vinker, E. Pajouheshgar, J. Y. Bo, R. C. Bachmann, A. H. Bermano, D. Cohen-Or, A. Zamir, and A. Shamir, “Clipasso: Semantically-aware object sketching,” arXiv , 2022
2022
Closest in time.
Y. Xu, H. Wei, M. Lin, Y. Deng, K. Sheng, M. Zhang, F. Tang, W. Dong, F. Huang, and C. Xu, “Transformers in computational visual media: A survey,” Computational Visual Media , 2022
2022
Closest in time.
F. Shamshad, S. Khan, S. W. Zamir, M. H. Khan, M. Hayat, F. S. Khan, and H. Fu, “Transformers in medical imaging: A survey,” arXiv , 2022
2022
Closest in time.
J. Selva, A. S. Johansen, S. Escalera, K. Nasrollahi, T. B. Moeslund, and A. Clapés, “Video transformers: A survey,” arXiv , 2022
2022
Closest in time.
F. Chen, D. Zhang, M. Han, X. Chen, J. Shi, S. Xu, and B. Xu, “Vlp: A survey on vision-language pre-training,” arXiv , 2022
2022
Closest in time.
F. Li, H. Zhang, Y.-F. Zhang, S. Liu, J. Guo, L. M. Ni, P. Zhang, and L. Zhang, “Vision-language intelligence: Tasks, representation learning, and large models,” arXiv , 2022
2022
Closest in time.
L. Yu, J. Chen, A. Sinha, M. M. Wang, H. Chen, T. L. Berg, and N. Zhang, “Commercemm: Large-scale commerce multimodal representation learning with omni retrieval,” arXiv , 2022
2022
Closest in time.
P. Xu, T. M. Hospedales, Q. Yin, Y.-Z. Song, T. Xiang, and L. Wang, “Deep learning for free-hand sketch: A survey,” TPAMI , 2022
2022
Closest in time.
Y.-L. Sung, J. Cho, and M. Bansal, “Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks,” in CVPR , 2022
2022
Closest in time.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” NeurIPS , 2022
2022
Closest in time.
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som et al. , “Image as a foreign language: Beit pretraining for all vision and vision-language tasks,” arXiv , 2022
2022
Closest in time.
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer et al. , “Pali: A jointly-scaled multilingual language-image model,” arXiv , 2022
2022
Closest in time.
S. Cao, P. Xu, and D. A. Clifton, “How to understand masked autoencoders,” arXiv , 2022
2022
Closest in time.
Z. Wang, N. Codella, Y.-C. Chen, L. Zhou, J. Yang, X. Dai, B. Xiao, H. You, S.-F. Chang, and L. Yuan, “Clip-td: Clip targeted distillation for vision-language tasks,” arXiv , 2022
2022
Closest in time.
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models,” arXiv , 2022
2022
Closest in time.
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” arXiv , 2022
2022
Closest in time.
J. Yang, J. Duan, S. Tran, Y. Xu, S. Chanda, L. Chen, B. Zeng, T. Chilimbi, and J. Huang, “Vision-language pre-training with triple contrastive learning,” arXiv , 2022
2022
Closest in time.
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang et al. , “Grounded language-image pre-training,” in CVPR , 2022
2022
Closest in time.
H. Zhang, P. Zhang, X. Hu, Y.-C. Chen, L. Li, X. Dai, L. Wang, L. Yuan, J.-N. Hwang, and J. Gao, “Glipv2: Unifying localization and vision-language understanding,” NeurIPS , 2022
2022
Closest in time.
T. Han, W. Xie, and A. Zisserman, “Temporal alignment networks for long-term video,” in CVPR , 2022
2022
Closest in time.
Y. Wang, X. Chen, L. Cao, W. Huang, F. Sun, and Y. Wang, “Multimodal token fusion for vision transformers,” in CVPR , 2022
2022
Closest in time.
Y. Wang, T. Ye, L. Cao, W. Huang, F. Sun, F. He, and D. Tao, “Bridged transformer for vision and point cloud 3d object detection,” in CVPR , 2022
2022
Closest in time.
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” in CVPR , 2022
2022
Closest in time.
N. Shvetsova, B. Chen, A. Rouditchenko, S. Thomas, B. Kingsbury, R. S. Feris, D. Harwath, J. Glass, and H. Kuehne, “Everything at once - multi-modal fusion transformer for video retrieval,” in CVPR , 2022
2022
Closest in time.
Z. Wang, Y. Wu, K. Narasimhan, and O. Russakovsky, “Multi-query video retrieval,” arXiv , 2022
2022
Closest in time.
V. Gabeur, A. Nagrani, C. Sun, K. Alahari, and C. Schmid, “Masking modalities for cross-modal video retrieval,” in WACV , 2022
2022
Closest in time.
S. Chen and B. Li, “Multi-modal dynamic graph transformer for visual grounding,” in CVPR , 2022
2022
Closest in time.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Tubedetr: Spatio-temporal video grounding with transformers,” in CVPR , 2022
2022
Closest in time.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” NeurIPS , 2022
2022
Closest in time.
S. Yan, X. Xiong, A. Arnab, Z. Lu, M. Zhang, C. Sun, and C. Schmid, “Multiview transformers for video recognition,” in CVPR , 2022
2022
Closest in time.
M. Ma, J. Ren, L. Zhao, D. Testuggine, and X. Peng, “Are multimodal transformers robust to missing modality?” in CVPR , 2022
2022
Closest in time.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” arXiv , 2022
2022
Closest in time.
R. Girdhar, M. Singh, N. Ravi, L. van der Maaten, A. Joulin, and I. Misra, “Omnivore: A single model for many visual modalities,” arXiv , 2022
2022
Closest in time.
T. Zhao, T. Zhang, M. Zhu, H. Shen, K. Lee, X. Lu, and J. Yin, “Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations,” arXiv , 2022
2022
Closest in time.
E. Aflalo, M. Du, S.-Y. Tseng, Y. Liu, C. Wu, N. Duan, and V. Lal, “Vl-interpret: An interactive visualization tool for interpreting vision-language transformers,” in CVPR , 2022
2022
Closest in time.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” arXiv , 2022
2022
Closest in time.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , 2023
2023
Closest in time.