Fetching the paper…
Reading the bibliography…
Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities.
Logik der Forschung
Karl Popper. 1935 · 1935
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Gool, Christopher K. Williams, John Winn, and Andrew Zisserman. 2010 · 2010
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Jize Cao, Zhe Gan, Yu Cheng, Licheng Yu, Yen-Chun Chen, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
VL-BERT: pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott. 2021 · 2021
Closest in time.
Measuring and increasing context usage in context-aware machine translation
Patrick Fernandes, Kayo Yin, Graham Neubig, and André F. T. Martins. 2021 · 2021
Closest in time.
Probing image-language transformers for verb understanding
Lisa Anne Hendricks and Aida Nematzadeh. 2021 · 2021
Closest in time.
MDETR – modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
What does BERT with vision look at?
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2020 · 2020
Cited alongside, same era.
12-in-1: Multi-task vision and language representation learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 2020 · 2020
Cited alongside, same era.
Stanza: A python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath R Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi. 2021 · 2021
Closest in time.
What context features can transformer language models use?
Joe O’Connor and Jacob Andreas. 2021 · 2021
Closest in time.
Seeing past words: Testing the cross-modal capabilities of pretrained v&l models on counting tasks
Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto. 2021 · 2021
Closest in time.
How much can clip benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. 2021 · 2021
Closest in time.