Revisiting visual question answering baselines
Original
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Original
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Later among the works it cites.
You only look once: Unified, real-time object detection
J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi · 2016
Later among the works it cites.
Zero-shot visual question answering
Original
D. Teney and A. van den Hengel · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola · 2016
Later among the works it cites.
Review networks for caption generation
Z. Yang, Y. Yuan, Y. Wu, R. Salakhutdinov, and W. W. Cohen · 2016
Later among the works it cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Closest in time.
Show, ask, attend, and answer: A strong baseline for visual question answering
Original
V. Kazemi and A. Elqursh · 2017
Closest in time.
Improved image captioning via policy gradient optimization of spider
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2017
Closest in time.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2017
Closest in time.
Areas of attention for image captioning
M. Pedersoli, T. Lucas, C. Schmid, and J. Verbeek · 2017
Closest in time.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2017
Closest in time.
Boosting image captioning with attributes
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei · 2017
Closest in time.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
D. Teney, P. Anderson, X. He, and A. van den Hengel · 2018
Closest in time.