Fetching the paper…
Reading the bibliography…
Attention networks in multimodal learning provide an efficient way to utilize given visual information selectively.
Modeling appearances with low-rank SVM
Lior Wolf, Hueihan Jhuang, and Tamir Hazan · 2007
Earlier work this paper cites.
Bilinear classifiers for visual recognition
Hamed Pirsiavash, Deva Ramanan, and Charless C. Fowlkes · 2009
Earlier work this paper cites.
Rectified Linear Units Improve Restricted Boltzmann Machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Multiscale combinatorial grouping
Pablo Arbeláez, Jordi Pont-Tuset, Jonathan T Barron, Ferran Marques, and Jitendra Malik · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Dropout : A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Edge boxes: Locating object proposals from edges
C Lawrence Zitnick and Piotr Dollár · 2014
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Spatial Transformer Networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Cited alongside, same era.
Multimodal Residual Learning for Visual QA
Jin-Hwa Kim, Sang-Woo Lee, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Hierarchical Question-Image Co-Attention for Visual Question Answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Dual Attention Networks for Multimodal Reasoning and Matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2016
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Later among the works it cites.
Query-Adaptive R-CNN for Open-Vocabulary Object Detection and Retrieval
Ryota Hinami and Shin’ichi Satoh · 2017
Later among the works it cites.
A Simple Loss Function for Improving the Convergence and Accuracy of Visual Question Answering Models
Ilija Ilievski and Jiashi Feng · 2017
Later among the works it cites.
Hadamard Product for Low-rank Bilinear Pooling
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2017
Later among the works it cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Cited alongside, same era.
Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
Tim Salimans and Diederik P. Kingma · 2016
Cited alongside, same era.
Residual Networks are Exponential Ensembles of Relatively Shallow Networks
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Cited alongside, same era.
Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering
Huijuan Xu and Kate Saenko · 2016
Cited alongside, same era.
Top-Down Neural Attention by Excitation Backprop
Jianming Zhang, Zhe Lin, Jonathan Brandt, Xiaohui Shen, and Stan Sclaroff · 2016
Cited alongside, same era.
Vqa: Visual question answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C Lawrence Zitnick, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2017
Cited alongside, same era.
YOLO9000: Better, Faster, Stronger
Joseph Redmon and Ali Farhadi · 2017
Later among the works it cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2017
Later among the works it cites.
Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge
Damien Teney, Peter Anderson, Xiaodong He, and Anton van den Hengel · 2017
Later among the works it cites.
The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
Peng Wang, Qi Wu, Chunhua Shen, and Anton van den Hengel · 2017
Later among the works it cites.
Interpretable and Globally Optimal Prediction for Textual Grounding using Image Concepts
Raymond A Yeh, Jinjun Xiong, Wen-Mei W Hwu, Minh N Do, and Alexander G Schwing · 2017
Later among the works it cites.
Interpretable Counting for Visual Question Answering
Alexander Trott, Caiming Xiong, and Richard Socher · 2018
Closest in time.
Beyond Bilinear: Generalized Multi-modal Factorized High-order Pooling for Visual Question Answering
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao · 2018
Closest in time.
Learning to Count Objects in Natural Images for Visual Question Answering
Yan Zhang, Jonathon Hare, and Adam Prügel-Bennett · 2018
Closest in time.