Fetching the paper…
Reading the bibliography…
Vision-Language (VL) models with the Two-Tower architecture have dominated visual-language representation learning in recent years.
Visual entailment: A novel task for fine-grained image understanding
Xie, N.; Lai, F.; Doran, D.; and Kadav, A. 2019 · 1901
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019b · 1907
Earlier work this paper cites.
VISUALBERT: ASimple AND PERFORMANT BASELINE FOR VISION AND LANGUAGE
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019 · 1908
Earlier work this paper cites.
A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks
Hashimoto, K.; Xiong, C.; Tsuruoka, Y.; and Socher, R. 2017 · 1933
Earlier work this paper cites.
Unifying Vision-and-Language Tasks via Text Generation
Cho, J.; Lei, J.; Tan, H.; and Bansal, M. 2021 · 1942
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Huang, Z.; Zeng, Z.; Liu, B.; Fu, D.; and Fu, J. 2020 · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; et al. 2009 · 2009
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. L. 2011 · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D.; and Fergus, R. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Li, F. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
Zhu, Y.; Kiros, R.; Zemel, R. S.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Ssd: Single shot multibox detector
Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; and Berg, A. C. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Sennrich, R.; Haddow, B.; and Birch, A. 2016 · 2016
Earlier work this paper cites.
Deep multi-task learning with low level tasks supervised at lower layers
Søgaard, A.; and Goldberg, Y. 2016 · 2016
Earlier work this paper cites.
What do Neural Machine Translation Models Learn about Morphology?
Belinkov, Y.; Durrani, N.; Dalvi, F.; Sajjad, H.; and Glass, J. 2017 · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Densely Connected Convolutional Networks
Huang, G.; Liu, Z.; van der Maaten, L.; and Weinberger, K. Q. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Feature Pyramid Networks for Object Detection
Lin, T.; Dollár, P.; Girshick, R. B.; He, K.; Hariharan, B.; and Belongie, S. J. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Exploiting Deep Representations for Neural Machine Translation
Dou, Z.-Y.; Tu, Z.; Wang, X.; Shi, S.; and Zhang, T. 2018 · 2018
Earlier work this paper cites.
Deep Contextualized Word Representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018a · 2018
Cited alongside, same era.
Dissecting Contextual Word Embeddings: Architecture and Representation
Peters, M. E.; Neumann, M.; Zettlemoyer, L.; and Yih, W.-t. 2018b · 2018
Cited alongside, same era.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Cited alongside, same era.
Dense Information Flow for Neural Machine Translation
Shen, Y.; Tan, X.; He, D.; Qin, T.; and Liu, T.-Y. 2018 · 2018
Cited alongside, same era.
Tips and Tricks for Visual Question Answering: Learnings From the 2017 Challenge
Teney, D.; Anderson, P.; He, X.; and van den Hengel, A. 2018 · 2018
Cited alongside, same era.
Multi-layer Representation Fusion for Neural Machine Translation
Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
Hendricks, L. A.; Mellor, J.; Schneider, R.; Alayrac, J.-B.; and Nematzadeh, A. 2021 · 2021
Later among the works it cites.
Unit: Multimodal multitask learning with a unified transformer
Hu, R.; and Singh, A. 2021 · 2021
Later among the works it cites.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Huang, Z.; Zeng, Z.; Huang, Y.; Liu, B.; Fu, D.; and Fu, J. 2021 · 2021
Later among the works it cites.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.; Parekh, Z.; Pham, H.; Le, Q. V.; Sung, Y.; Li, Z.; and Duerig, T. 2021 · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, Q.; Li, F.; Xiao, T.; Li, Y.; Li, Y.; and Zhu, J. 2018 · 2018
Cited alongside, same era.
Deep Layer Aggregation
Yu, F.; Wang, D.; Shelhamer, E.; and Darrell, T. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement
Dou, Z.; Tu, Z.; Wang, X.; Wang, L.; Shi, S.; and Zhang, T. 2019 · 2019
Cited alongside, same era.
What Does BERT Learn about the Structure of Language?
Jawahar, G.; Sagot, B.; and Seddah, D. 2019 · 2019
Cited alongside, same era.
Panoptic Feature Pyramid Networks
Kirillov, A.; Girshick, R. B.; He, K.; and Dollár, P. 2019 · 2019
Cited alongside, same era.
Linguistic Knowledge and Transferability of Contextual Representations
Liu, N. F.; Gardner, M.; Belinkov, Y.; Peters, M. E.; and Smith, N. A. 2019a · 2019
Cited alongside, same era.
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Later among the works it cites.
Attention bottlenecks for multimodal fusion
Nagrani, A.; Yang, S.; Arnab, A.; Jansen, A.; Schmid, C.; and Sun, C. 2021 · 2021
Later among the works it cites.
Intriguing properties of vision transformers
Naseer, M. M.; Ranasinghe, K.; Khan, S. H.; Hayat, M.; Shahbaz Khan, F.; and Yang, M.-H. 2021 · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Ni, M.; Huang, H.; Su, L.; Cui, E.; Bharti, T.; Wang, L.; Zhang, D.; and Duan, N. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Raghu, M.; Unterthiner, T.; Kornblith, S.; Zhang, C.; and Dosovitskiy, A. 2021 · 2021
Later among the works it cites.
How Much Can CLIP Benefit Vision-and-Language Tasks?
Shen, S.; Li, L. H.; Tan, H.; Bansal, M.; Rohrbach, A.; Chang, K.-W.; Yao, Z.; and Keutzer, K. 2021 · 2021
Later among the works it cites.
Visual Grounding Strategies for Text-Only Natural Language Processing
Sileo, D. 2021 · 2021
Later among the works it cites.
FLAVA: A Foundational Language And Vision Alignment Model
Singh, A.; Hu, R.; Goswami, V.; Couairon, G.; Galuba, W.; Rohrbach, M.; and Kiela, D. 2021 · 2021
Later among the works it cites.
Xgpt: Cross-modal generative pre-training for image captioning
Xia, Q.; Huang, H.; Duan, N.; Zhang, D.; Ji, L.; Sui, Z.; Cui, E.; Bharti, T.; and Zhou, M. 2021 · 2021
Later among the works it cites.
SegFormer: Simple and efficient design for semantic segmentation with transformers
Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; and Luo, P. 2021 · 2021
Later among the works it cites.
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Zeng, Y.; Zhang, X.; and Li, H. 2021 · 2021
Later among the works it cites.
LiT: Zero-Shot Transfer with Locked-image Text Tuning
Zhai, X.; Wang, X.; Mustafa, B.; Steiner, A.; Keysers, D.; Kolesnikov, A.; and Beyer, L. 2021 · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P. H.; et al. 2021 · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Closest in time.
Pali: A jointly-scaled multilingual language-image model
Chen, X.; Wang, X.; Changpinyo, S.; Piergiovanni, A.; Padlewski, P.; Salz, D.; Goodman, S.; Grycner, A.; Mustafa, B.; Beyer, L.; et al. 2022 · 2022
Closest in time.
An Empirical Study of Training End-to-End Vision-and-Language Transformers
Dou, Z.-Y.; Xu, Y.; Gan, Z.; Wang, J.; Wang, S.; Wang, L.; Zhu, C.; Zhang, P.; Yuan, L.; Peng, N.; Liu, Z.; and Zeng, M. 2022 · 2022
Closest in time.
A Survey of Vision-Language Pre-Trained Models
Du, Y.; Liu, Z.; Li, J.; and Zhao, W. X. 2022 · 2022
Closest in time.
UNIMO-2: End-to-End Unified Vision-Language Grounded Learning
Li, W.; Gao, C.; Niu, G.; Xiao, X.; Liu, H.; Liu, J.; Wu, H.; and Wang, H. 2022c · 2022
Closest in time.
KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation
Liu, Y.; Wu, C.; Tseng, S.-Y.; Lal, V.; He, X.; and Duan, N. 2022 · 2022
Closest in time.
Revealing the Dark Secrets of Masked Image Modeling
Xie, Z.; Geng, Z.; Hu, J.; Zhang, Z.; Hu, H.; and Cao, Y. 2022 · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Yu, J.; Wang, Z.; Vasudevan, V.; Yeung, L.; Seyedhosseini, M.; and Wu, Y. 2022 · 2022
Closest in time.
Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision
Tan, H.; and Bansal, M. 2020 · 2080
Closest in time.