Fetching the paper…
Reading the bibliography…
Despite significant progress in Visual Question Answering over the years, robustness of today's VQA models leave much to be desired.
Robust processing of real-world natural-language texts
J. R. Hobbs, D. E. Appelt, J. Bear, and M. Tyson · 1992
Earlier work this paper cites.
The search for robustness in natural language understanding
M. Stede · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Automation of question generation from sentences
H. Ali, Y. Chali, and S. A. Hasan · 2010
Earlier work this paper cites.
Natural language question generation using syntax and keywords
S. Kalady, A. Elikkottil, and R. Das · 2010
Earlier work this paper cites.
Dense point trajectories by gpu-accelerated large displacement optical flow
N. Sundaram, T. Brox, and K. Keutzer · 2010
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Earlier work this paper cites.
Skip-thought vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
A. Agrawal, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Dual learning for machine translation
D. He, Y. Xia, T. Qin, L. Wang, N. Yu, T. Liu, and W.-Y. Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Generating natural questions about an image
N. Mostafazadeh, I. Misra, J. Devlin, M. Mitchell, X. He, and L. Vanderwende · 2016
Cited alongside, same era.
Generating factoid questions with recurrent neural networks: The 30m factoid question-answer corpus
I. V. Serban, A. García-Durán, C. Gulcehre, S. Ahn, S. Chandar, A. Courville, and Y. Bengio · 2016
Cited alongside, same era.
M. Spranger, J. Suchan, and M. Bhatt · 2016
Unpaired image-to-image translation using cycle-consistent adversarial networks
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros · 2017
Later among the works it cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi · 2018
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Detectron
R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He · 2018
Later among the works it cites.
Adversarial example generation with syntactically controlled paraphrase networks
M. Iyyer, J. Wieting, K. Gimpel, and L. Zettlemoyer · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering
H. Ben-Younes, R. Cadene, M. Cord, and N. Thome · 2017
Cited alongside, same era.
Towards linguistically generalizable nlp systems: A workshop and shared task
A. Ettinger, S. Rao, H. Daumé III, and E. M. Bender · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Cited alongside, same era.
Creativity: Generating diverse questions using variational autoencoders
U. Jain, Z. Zhang, and A. Schwing · 2017
Cited alongside, same era.
J.-H. Kim, J. Jun, and B.-T. Zhang · 2018
Later among the works it cites.
Visual question generation as dual task of visual question answering
Y. Li, N. Duan, B. Zhou, X. Chu, W. Ouyang, X. Wang, and M. Zhou · 2018
Later among the works it cites.
ivqa: Inverse visual question answering
F. Liu, T. Xiang, T. M. Hospedales, W. Yang, and C. Sun · 2018
Later among the works it cites.
The natural language decathlon: Multitask learning as question answering, 2018
B. McCann, N. S. Keskar, C. Xiong, and R. Socher · 2018
Later among the works it cites.
Learning by asking questions
I. Misra, R. Girshick, R. Fergus, M. Hebert, A. Gupta, and L. van der Maaten · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Later among the works it cites.
Learning to collaborate for question answering and asking
D. Tang, N. Duan, Z. Yan, Z. Zhang, Y. Sun, S. Liu, Y. Lv, and M. Zhou · 2018
Later among the works it cites.
Qg-net: A data-driven question generation model for educational content
Z. Wang, A. S. Lan, W. Nie, A. E. Waters, P. J. Grimaldi, and R. G. Baraniuk · 2018
Later among the works it cites.
Fooling vision and language models despite localization and attention mechanism
X. Xu, X. Chen, C. Liu, A. Rohrbach, T. Darrell, and D. Song · 2018
Later among the works it cites.
Pythia v0.1: the winning entry to the vqa challenge 2018
Yu Jiang*, Vivek Natarajan*, Xinlei Chen*, M. Rohrbach, D. Batra, and D. Parikh · 2018
Later among the works it cites.