Fetching the paper…
Reading the bibliography…
We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images.
Yarbus, A. L · 1967
Earlier work this paper cites.
The dynamic representation of scenes
Rensink, Ronald A · 2000
Earlier work this paper cites.
What do we perceive in a glance of a real-world scene?
Fei-Fei, Li, Iyer, Asha, Koch, Christof, and Perona, Pietro · 2007
Earlier work this paper cites.
Learning to predict where humans look
Judd, Tilke, Ehinger, Krista, Durand, Frédo, and Torralba, Antonio · 2009
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Saliency in Crowd
Jiang, Ming, Xu, Juan, and Zhao, Qi · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context, 2014
Lin, Tsung-Yi, Maire, Michael, Belongie, Serge, Hays, James, Perona, Pietro, Ramanan, Deva, Dollár, Piotr, and Zitnick, C. Lawrence · 2014
Earlier work this paper cites.
Recurrent Models of Visual Attention
Mnih, Volodymyr, Heess, Nicolas, Graves, Alex, and Kavukcuoglu, Koray · 2014
Cited alongside, same era.
Attention for fine-grained categorization
Sermanet, Pierre, Frome, Andrea, and Real, Esteban · 2014
Cited alongside, same era.
Vqa: Visual question answering
Antol, Stanislaw, Agrawal, Aishwarya, Lu, Jiasen, Mitchell, Margaret, Batra, Dhruv, Zitnick, C. Lawrence, and Parikh, Devi · 2015
Cited alongside, same era.
Multiple Object Recognition With Visual Attention
Ba, Jimmy Lei, Mnih, Volodymyr, and Kavukcuoglu, Koray · 2015
Cited alongside, same era.
Describing multimedia content using attention-based encoder-decoder networks
Cho, KyungHyun, Courville, Aaron C., and Bengio, Yoshua · 2015
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Xu, Huijuan and Saenko, Kate · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron C., Salakhutdinov, Ruslan, Zemel, Richard S., and Bengio, Yoshua · 2015
Later among the works it cites.
Stacked attention networks for image question answering
Yang, Zichao, He, Xiaodong, Gao, Jianfeng, Deng, Li, and Smola, Alexander J · 2015
Later among the works it cites.
Learning to compose neural networks for question answering
Andreas, Jacob, Rohrbach, Marcus, Darrell, Trevor, and Klein, Dan · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, Jacob, Gupta, Saurabh, Girshick, Ross, Mitchell, Margaret, and Zitnick, C. Lawrence · 2015
Cited alongside, same era.
Salicon: Saliency in context
Jiang, Ming, Huang, Shengsheng, Duan, Juanyong, and Zhao, Qi · 2015
Cited alongside, same era.
Firat, Orhan, Cho, KyungHyun, and Bengio, Yoshua · 2016
Closest in time.
Hierarchical Co-Attention for Visual Question Answering
Lu, J., Yang, J., Batra, D., and Parikh, D · 2016
Closest in time.
Dynamic memory networks for visual and textual question answering
Xiong, Caiming, Merity, Stephen, and Socher, Richard · 2016
Closest in time.