Pythia v0. 1: the winning entry to the vqa challenge 2018
Original
Yu Jiang, Vivek Natarajan, Xinlei Chen, Marcus Rohrbach, Dhruv Batra, and Devi Parikh. 2018 · 2018
Later among the works it cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018 · 2018
Later among the works it cites.
Out of the box: Reasoning with graph convolution nets for factual visual question answering
Medhini Narasimhan, Svetlana Lazebnik, and Alexander Schwing. 2018 · 2018
Later among the works it cites.
Straight to the facts: Learning knowledge base retrieval for factual visual question answering
Medhini Narasimhan and Alexander G Schwing. 2018 · 2018
Later among the works it cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Later among the works it cites.
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton van den Hengel. 2018 · 2018
Later among the works it cites.
Deep cross-modal projection learning for image-text matching
Ying Zhang and Huchuan Lu. 2018 · 2018
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Later among the works it cites.
Camp: Cross-modal adaptive message passing for text-image retrieval
Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao. 2019 · 2019
Later among the works it cites.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Multi-modality cross attention network for image and sentence matching
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu. 2020 · 2020
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Original
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Closest in time.
Visualsparta: Sparse transformer fragment-level matching for large-scale text-to-image search
Original
Xiaopeng Lu, Tiancheng Zhao, and Kyusong Lee. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Original
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Closest in time.
Learning fragment self-attention embeddings for image-text matching
Yiling Wu, Shuhui Wang, Guoli Song, and Qingming Huang. 2019 · 2096
Closest in time.