Fetching the paper…
Reading the bibliography…
In this work, we address multi-modal information needs that contain text questions and images by focusing on passage retrieval for outside-knowledge visual question answering.
Combination of Multiple Searches
E. A. Fox and J. A. Shaw · 1993
Earlier work this paper cites.
Analyses of Multiple Evidence Combination
J. H. Lee · 1997
Earlier work this paper cites.
The TREC-8 Question Answering Track Evaluation
E. M. Voorhees and D. M. Tice · 1999
Earlier work this paper cites.
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
L. Xiong, C. Xiong, Y. Li, K.-F. Tang, J. Liu, P. Bennett, J. Ahmed, and A. Overwijk · 2007
Earlier work this paper cites.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
G. V. Cormack, C. L. A. Clarke, and S. Büttcher · 2009
Earlier work this paper cites.
Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
W. Xiong, X. Li, S. Iyer, J. Du, P. Lewis, W. Y. Wang, Y. Mehdad, W. tau Yih, S. Riedel, D. Kiela, and B. Ouguz · 2009
Earlier work this paper cites.
Y. Qu, Y. Ding, J. Liu, K. Liu, R. Ren, X. Zhao, D. Dong, H. Wu, and H. Wang · 2010
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Parikh, and D. Batra · 2015
Earlier work this paper cites.
Ask Your Neurons: A Neural-Based Approach to Answering Questions about Images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
S. Ren, K. He, R. B. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Visual Madlibs: Fill in the Blank Description Generation and Question Answering
L. Yu, E. Park, A. Berg, and T. Berg · 2015
Earlier work this paper cites.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
Hierarchical Question-Image Co-Attention for Visual Question Answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Ask Me Anything: Free-Form Visual Question Answering Based on Knowledge from External Sources
Q. Wu, P. Wang, C. Shen, A. Dick, and A. V. D. Hengel · 2016
Cited alongside, same era.
Dynamic Memory Networks for Visual and Textual Question Answering
C. Xiong, S. Merity, and R. Socher · 2016
Cited alongside, same era.
Visual7W: Grounded Question Answering in Images
Y. Zhu, O. Groth, M. S. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
MUTAN: Multimodal Tucker Fusion for Visual Question Answering
H. Ben-younes, R. Cadène, M. Cord, and N. Thome · 2017
Cited alongside, same era.
Risk-Reward Trade-offs in Rank Fusion
R. Benham and J. S. Culpepper · 2017
Cited alongside, same era.
Reading Wikipedia to Answer Open-Domain Questions
D. Chen, A. Fisch, J. Weston, and A. Bordes · 2017
Cited alongside, same era.
FVQA: Fact-Based Visual Question Answering
P. Wang, Q. Wu, C. Shen, A. Dick, and A. van den Hengel · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Later among the works it cites.
Latent Retrieval for Weakly Supervised Open Domain Question Answering
K. Lee, M.-W. Chang, and K. Toutanova · 2019
Later among the works it cites.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Later among the works it cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
H. H. Tan and M. Bansal · 2019
Later among the works it cites.
ConceptBert: Concept-Aware Representation for Visual Question Answering
F. Gardères, M. Ziaeefard, B. Abeloos, and F. Lécué · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Billion-scale similarity search with GPUs
J. Johnson, M. Douze, and H. Jégou · 2017
Cited alongside, same era.
Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks
G. Li, H. Su, and W. Zhu · 2017
Cited alongside, same era.
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Explicit Knowledge-based Reasoning for Visual Question Answering
P. Wang, Q. Wu, C. Shen, A. Dick, and A. V. D. Hengel · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Cited alongside, same era.
Later among the works it cites.
REALM: Retrieval-Augmented Language Model Pre-Training
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang · 2020
Later among the works it cites.
Dense Passage Retrieval for Open-Domain Question Answering
V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Y. Wu, S. Edunov, D. Chen, and W. tau Yih · 2020
Later among the works it cites.
Recipe Retrieval with Visual Query of Ingredients
Y.-C. Lien, H. Zamani, and W. B. Croft · 2020
Later among the works it cites.
Sparse, Dense, and Attentional Representations for Text Retrieval
Y. Luan, J. Eisenstein, K. Toutanova, and M. Collins · 2020
Later among the works it cites.
IART: Intent-Aware Response Ranking with Transformers in Information-Seeking Conversation Systems
L. Yang, M. Qiu, C. Qu, C. Chen, J. Guo, Y. Zhang, W. B. Croft, and H. Chen · 2020
Later among the works it cites.
Cross-modal Knowledge Reasoning for Knowledge-based Visual Question Answering
J. Yu, Z. Zhu, Y. Wang, W. Zhang, Y. Hu, and J. Tan · 2020
Later among the works it cites.
Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering
Z. Zhu, J. Yu, Y. Wang, Y. Sun, Y. Hu, and Q. Wu · 2020
Later among the works it cites.
Towards Multi-Modal Conversational Information Seeking
Y. Deldjoo, J. R. Trippas, and H. Zamani · 2021
Closest in time.
Weakly-Supervised Open-Retrieval Conversational Question Answering
C. Qu, L. Yang, C. Chen, W. Croft, K. Krishna, and M. Iyyer · 2021
Closest in time.