Fetching the paper…
Reading the bibliography…
Large vision-and-language models (VLMs) trained to match images with text on large-scale datasets of image-text pairs have shown impressive generalization ability on several vision and language tasks.
Wordnet: a lexical database for english
G. A. Miller · 1995
Earlier work this paper cites.
Nltk: the natural language toolkit
S. Bird · 2006
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Earlier work this paper cites.
Vinvl: Revisiting visual representations in vision-language models
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al · 2020
Cited alongside, same era.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
H. Chefer, S. Gur, and L. Wolf · 2021
Cited alongside, same era.
Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers
S. Frank, E. Bugliarello, and D. Elliott · 2021
Cited alongside, same era.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Y. Zeng, X. Zhang, and H. Li · 2021
Later among the works it cites.
Towards general purpose vision systems: An end-to-end task-agnostic vision-language architecture
T. Gupta, A. Kamath, A. Kembhavi, and D. Hoiem · 2022
Later among the works it cites.
F. Liu, G. Emerson, and N. Collier · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela · 2022
Later among the works it cites.
Reclip: A strong zero-shot baseline for referring expression comprehension
S. Subramanian, W. Merrill, T. Darrell, M. Gardner, S. Singh, and A. Rohrbach · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. A. Hendricks and A. Nematzadeh · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Simvlm: Simple visual language model pretraining with weak supervision
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao · 2021
Cited alongside, same era.
What’s" up" with vision-language models? investigating their struggle with spatial reasoning
A. Kamath, J. Hessel, and K.-W. Chang
Cited in the paper.
Later among the works it cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Later among the works it cites.
When and why vision-language models behave like bag-of-words models, and what to do about it?
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou · 2022
Later among the works it cites.
Weakly-supervised learning of visual relations in multimodal pretraining
E. Bugliarello, A. Nematzadeh, and L. A. Hendricks · 2023
Closest in time.
Incorporating structured representations into pretrained vision & language models using scene graphs
R. Herzig, A. Mendelson, L. Karlinsky, A. Arbelle, R. Feris, T. Darrell, and A. Globerson · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.