Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) is an emerging area of interest for researches, being a recent problem in natural language processing and image prediction.
Estimating the dimension of a model
Gideon Schwarz · 1978
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
The expectation-maximization algorithm
T.K. Moon · 1996
Earlier work this paper cites.
The jensen-shannon divergence
ML Menéndez, JA Pardo, L Pardo, and MC Pardo · 1997
Earlier work this paper cites.
Child: A first step towards continual learning
Mark B. Ring · 1997
Earlier work this paper cites.
Stochastic complexity in statistical inquiry, 1998
J Rissanen · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard Hovy · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Chin-Yew Lin and Franz Josef Och · 2004
Earlier work this paper cites.
Understanding inverse document frequency: on theoretical arguments for idf
Stephen Robertson · 2004
Earlier work this paper cites.
D³ data-driven documents
Michael Bostock, Vadim Ogievetsky, and Jeffrey Heer · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Word spotting and recognition with embedded attributes
Jon Almazán, Albert Gordo, Alicia Fornés, and Ernest Valveny · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2017
Earlier work this paper cites.
Enriching Word Vectors with Subword Information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
FigureQA: An Annotated Figure Dataset for Visual Reasoning
Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Akos Kadar, Adam Trischler, and Yoshua Bengio · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Earlier work this paper cites.
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Earlier work this paper cites.
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton van den Hengel · 2018
Earlier work this paper cites.
Rosetta: Large scale system for text detection and recognition in images
Fedor Borisyuk, Albert Gordo, and Viswanath Sivakumar · 2018
Earlier work this paper cites.
BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages
Benjamin Heinzerling and Michael Strube · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
K. Kafle, B. Price, S. Cohen, and C. Kanan · 2018
Earlier work this paper cites.
Attention-based deep multiple instance learning for visual concept detection
Chao Zhang, Yanwei Pang, Jia Zhu, Cunzhao Shi, Jian Zhang, and Chunfeng Yuan · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
Damien Teney, Peter Anderson, Xiaodong He, and Anton van den Hengel · 2018
Cited alongside, same era.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Cited alongside, same era.
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marcal Rusinol, Ernest Valveny, C.V. Jawahar, and Dimosthenis Karatzas · 2019
Cited alongside, same era.
Multimodal continuous visual attention mechanisms
A. Farinhas, A. T. Martins, and P. Q. Aguiar · 2021
Later among the works it cites.
Roses are red, violets are blue… but should vqa expect them to?
Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf · 2021
Later among the works it cites.
Unshuffling data for improved generalization in visual question answering
D. Teney, E. Abbasnejad, and A. van den Hengel · 2021
Later among the works it cites.
Structured multimodal attentions for textvqa
Chenyu Gao, Qi Zhu, Peng Wang, Hui Li, Yuliang Liu, Anton Van den Hengel, and Qi Wu · 2021
Later among the works it cites.
Zero-shot visual question answering using knowledge graph
Zhuo Chen, Jiaoyan Chen, Yuxia Geng, Jeff Z Pan, Zonggang Yuan, and Huajun Chen · 2021
Later among the works it cites.
Iq-vqa: Intelligent visual question answering
Vatsal Goel, Mohit Chandak, Ashish Anand, and Prithwijit Guha · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Claudio Greco, Barbara Plank, Raquel Fernández, and Raffaella Bernardi · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2019
Cited alongside, same era.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Cited alongside, same era.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Later among the works it cites.
Separating skills and concepts for novel visual question answering
S. Whitehead, H. Wu, H. Ji, R. Feris, and K. Saenko · 2021
Later among the works it cites.
Debiased visual question answering from feature and sample perspectives
Zhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu, and Qi Wu · 2021
Later among the works it cites.
Graphhopper: Multi-hop scene graph reasoning for visual question answering
Rajat Koner, Hang Li, Marcel Hildebrandt, Deepan Das, Volker Tresp, and Stephan Günnemann · 2021
Later among the works it cites.
Mirtt: Learning multimodal interaction representations from trilinear transformers for visual question answering
Junjie Wang, Yatai Ji, Jiaqi Sun, Yujiu Yang, and Tetsuya Sakai · 2021
Later among the works it cites.
Contrastive Pre-training and Representation Distillation for Medical Visual Question Answering Based on Radiology Images
Bo Liu, Li-Ming Zhan, and Xiao-Ming Wu · 2021
Later among the works it cites.
Question-controlled text-aware image captioning
Anwen Hu, Shizhe Chen, and Qin Jin · 2021
Later among the works it cites.
Sysu-hcp at vqa-med 2021: A data-centric model with efficient training methodology for medical visual question answering
Haifan Gong, Ricong Huang, Guanqi Chen, and Guanbin Li · 2021
Later among the works it cites.
Greedy gradient ensemble for robust visual question answering
Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, and Qi Tian · 2021
Later among the works it cites.
LPF: A language-prior feedback objective function for de-biased visual question answering
Zujie Liang, Haifeng Hu, and Jiaying Zhu · 2021
Later among the works it cites.
Passage retrieval for outside-knowledge visual question answering
Chen Qu, Hamed Zamani, Liu Yang, W. Bruce Croft, and Erik Learned-Miller · 2021
Later among the works it cites.
How transferable are reasoning patterns in vqa?
C. Kervadec, T. Jaunet, G. Antipov, M. Baccouche, R. Vuillemot, and C. Wolf · 2021
Later among the works it cites.
Mdetr - modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Later among the works it cites.
Tap: Text-aware pre-training for text-vqa and text-caption
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo · 2021
Later among the works it cites.
Counterfactual VQA: A cause-effect look at language bias
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Later among the works it cites.
Overview of the ImageCLEF 2021: Multimedia retrieval in medical, nature, internet and social media applications
Bogdan Ionescu, Henning Müller, Renaud Péteri, Asma Ben Abacha, Mourad Sarrouti, Dina Demner-Fushman, Sadid A. Hasan, Serge Kozlovski, Vitali Liauchuk, Yashin Dicente, Vassili Kovalev, Obioma Pelka, Alba García Seco de Herrera, Janadhip Jacutprakart, Christoph M. Friedrich, Raul Berari, Andrei Tauteanu, Dimitri Fichou, Paul Brie, Mihai Dogariu, Liviu Daniel Ştefan, Mihai Gabriel Constantin, Jon Chamberlain, Antonio Campello, Adrian Clark, Thomas A. Oliver, Hassan Moustahfid, Adrian Popescu, and Jérôme Deshayes-Chossart · 2021
Later among the works it cites.
VisQA: X-raying Vision and Language Reasoning in Transformers
Theo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov, Moez Baccouche, and Christian Wolf · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Bringing light into the dark: A large-scale evaluation of knowledge graph embedding models under a unified framework
Mehdi Ali, Max Berrendorf, Charles Hoyt, Laurent Vermue, Mikhail Galkin, Sahand Sharifzadeh, Asja Fischer, Volker Tresp, and Jens Lehmann · 2021
Later among the works it cites.
Medical visual question answering: A survey, 2022
Zhihong Lin, Donghao Zhang, Qingyi Tac, Danli Shi, Gholamreza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge · 2022
Later among the works it cites.
Coarse-to-fine reasoning for visual question answering
Binh X Nguyen, Tuong Do, Huy Tran, Erman Tjiputra, Quang D Tran, and Anh Nguyen · 2022
Later among the works it cites.
Answer questions with right image regions: A visual attention regularization approach
Yibing Liu, Yangyang Guo, Jianhua Yin, Xuemeng Song, Weifeng Liu, Liqiang Nie, and Min Zhang · 2022
Later among the works it cites.
CLEVR-x: A visual reasoning dataset for natural language explanations
Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch, and Zeynep Akata · 2022
Later among the works it cites.
VLMo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei · 2022
Later among the works it cites.
Latr: Layout-aware transformer for scene-text vqa
Ali Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju, and R. Manmatha · 2022
Later among the works it cites.
Weakly supervised grounding for vqa in vision-language transformers
Aisha Urooj Khan, Hilde Kuehne, Chuang Gan, Niels da Vitoria Lobo, and Mubarak Shah · 2022
Later among the works it cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks, 2022
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei · 2022
Later among the works it cites.
Must-vqa: Multilingual scene-text vqa
Emanuele Vivoli, Ali Furkan Biten, Andres Mafla, Dimosthenis Karatzas, and Lluis Gomez · 2023
Closest in time.
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Alexander Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut · 2023
Closest in time.