Fetching the paper…
Reading the bibliography…
Multimodal IR, spanning text corpus, knowledge graph and images, called outside knowledge visual question answering (OKVQA), is of much recent interest.
Dynamically Fused Graph Network for Multi-hop Reasoning
Yunxuan Xiao, Yanru Qu, Lin Qiu, Hao Zhou, Lei Li, Weinan Zhang, and Yong Yu. 2019 · 1905
Earlier work this paper cites.
Introduction to WordNet: an on-line lexical database
G. A. Miller, R. Beckwith, C. Fellbaum, D. Gross, and K. J. Miller. 1990 · 1990
Earlier work this paper cites.
The TREC-8 Question Answering Track Report. In TREC
Ellen M Voorhees. 1999 · 1999
Earlier work this paper cites.
Overview of the TREC 2001 Question Answering Track. In The Tenth Text REtrieval Conference (NIST Special Publication) , Vol. 500-250. 42–51
Ellen Voorhees. 2001 · 2001
Earlier work this paper cites.
Do Multi-Hop Question Answering Systems Know How to Answer the Single-Hop Sub-Questions?
Yixuan Tang, Hwee Tou Ng, and Anthony KH Tung. 2020 · 2002
Earlier work this paper cites.
Term proximity scoring for ad-hoc retrieval on very large text collections. In SIGIR Conference (Seattle, Washington, USA). ACM, 621–622
Stefan Büttcher, Charles L. A. Clarke, and Brad Lushman. 2006 · 2006
Earlier work this paper cites.
Proximity-based document representation for named entity retrieval. In CIKM . ACM, 731–740
Desislava Petkova and W Bruce Croft. 2007 · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge. In SIGMOD Conference . 1247–1250
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008 · 2008
Earlier work this paper cites.
Positional language models for information retrieval. In SIGIR Conference . 299–306
Yuanhua Lv and ChengXiang Zhai. 2009 · 2009
Earlier work this paper cites.
Cross-modal knowledge reasoning for knowledge-based visual question answering
Jing Yu, Zihao Zhu, Yujing Wang, Weifeng Zhang, Yue Hu, and Jianlong Tan. 2020 · 2009
Earlier work this paper cites.
Open Information Extraction: The Second Generation. In IJCAI . 3–10
Oren Etzioni, Anthony Fader, Janara Christensen, Stephen Soderland, and Mausam Mausam. 2011 · 2011
Earlier work this paper cites.
Semantic Parsing on Freebase from Question-Answer Pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Seattle, Washington, USA, 1533–1544
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013b · 2013
Earlier work this paper cites.
Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014, Columbus, OH, USA, June 23-28, 2014 . IEEE Computer Society, 580–587
Ross B. Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Zitnick. 2014 · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander Berg, and Li Fei-Fei. 2014 · 2014
Earlier work this paper cites.
Information Extraction over Structured Data: Question Answering with Freebase. In ACL Conference . ACL
Xuchen Yao and Benjamin Van Durme. 2014 · 2014
Cited alongside, same era.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada , Corinna Cortes, Neil D. Lawrence, Daniel D. Lee, Masashi Sugiyama, and Roman Garnett (Eds.). 91–99
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Austin, Texas, 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
GQA: A new dataset for real-world visual reasoning and compositional question answering. In CVPR . 6700–6709
Drew A Hudson and Christopher D Manning. 2019 · 2019
Later among the works it cites.
OK-VQA: A visual question answering benchmark requiring external knowledge. In CVPR . 3195–3204
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Later among the works it cites.
Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 3428–3448
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Compositional Questions Do Not Necessitate Multi-hop Reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 4249–4257
Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering. In Proceedings of the IEEE international conference on computer vision . 2612–2620
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Cited alongside, same era.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Poincaré embeddings for learning hierarchical representations. In Advances in neural information processing systems . 6338–6347
Maximillian Nickel and Douwe Kiela. 2017 · 2017
Cited alongside, same era.
ConceptNet 5.5: An Open Multilingual Graph of General Knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi. 2017 · 2017
Cited alongside, same era.
Think you have solved question answering? try ARC, the AI2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Multimodal explanations: Justifying decisions and pointing to the evidence. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 8779–8788
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2018 · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 2019
Later among the works it cites.
Are Red Roses Red? Evaluating Consistency of Question-Answering Models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 6174–6184
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Later among the works it cites.
PullNet: Open Domain Question Answering with Iterative Retrieval on Knowledge Bases and Text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . Association for Computational Linguistics, Hong Kong, China, 2380–2390
Haitian Sun, Tania Bedrax-Weiss, and William Cohen. 2019 · 2019
Later among the works it cites.
MultiQA: An Empirical Investigation of Generalization and Transfer in Reading Comprehension. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 4911–4921
Alon Talmor and Jonathan Berant. 2019 · 2019
Later among the works it cites.
Complex Question Decomposition for Semantic Parsing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 4477–4486
Haoyu Zhang, Jingjing Cai, Jianjun Xu, and Ji Wang. 2019 · 2019
Later among the works it cites.
ConceptBert: Concept-Aware Representation for Visual Question Answering. In Findings of the Association for Computational Linguistics: EMNLP 2020 . Association for Computational Linguistics, Online, 489–498
François Gardères, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue. 2020 · 2020
Later among the works it cites.
SpanBERT: Improving Pre-training by Representing and Predicting Spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Later among the works it cites.
The Open Images Dataset V4: Unified Image Classification, Object Detection, and Visual Relationship Detection at Scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. 2020 · 2020
Later among the works it cites.
Unsupervised Question Decomposition for Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, Online, 8864–8880
Ethan Perez, Patrick Lewis, Wen-tau Yih, Kyunghyun Cho, and Douwe Kiela. 2020 · 2020
Later among the works it cites.
How Much Knowledge Can You Pack into the Parameters of a Language Model?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 5418–5426
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Later among the works it cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning . PMLR, 11328–11339
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020 · 2020
Later among the works it cites.
Mediawiki Parser
Wikimedia Foundation. 2021 · 2021
Closest in time.