Fetching the paper…
Reading the bibliography…
We propose an end-to-end-trainable attention module for convolutional neural network (CNN) architectures built for image classification.
One-shot learning of object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2006
Earlier work this paper cites.
Caltech-256 object category dataset
Gregory Griffin, Alex Holub, and Pietro Perona · 2007
Earlier work this paper cites.
What, where and who? classifying events by scene and object recognition
Li-Jia Li and Li Fei-Fei · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Recognizing indoor scenes
Ariadna Quattoni and Antonio Torralba · 2009
Earlier work this paper cites.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning [http://cs.stanford.edu/ acoates/stl10]
Adam Coates, Honglak Lee, and Andrew Y Ng · 2010
Earlier work this paper cites.
Discriminative clustering for image co-segmentation
Armand Joulin, Francis Bach, and Jean Ponce · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Human action recognition by learning bases of action attributes and parts
Bangpeng Yao, Xiaoye Jiang, Aditya Khosla, Andy Lai Lin, Leonidas Guibas, and Li Fei-Fei · 2011
Earlier work this paper cites.
Multi-class cosegmentation
Armand Joulin, Francis Bach, and Jean Ponce · 2012
Earlier work this paper cites.
Saliency detection via absorbing markov chain
Bowen Jiang, Lihe Zhang, Huchuan Lu, Chuan Yang, and Ming-Hsuan Yang · 2013
Earlier work this paper cites.
Unsupervised joint object discovery and segmentation in internet images
Michael Rubinstein, Armand Joulin, Johannes Kopf, and Ce Liu · 2013
Cited alongside, same era.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Cited alongside, same era.
Saliency detection: A boolean map approach
Jianming Zhang and Stan Sclaroff · 2013
Cited alongside, same era.
Multiscale combinatorial grouping
Pablo Arbeláez, Jordi Pont-Tuset, Jonathan T Barron, Ferran Marques, and Jitendra Malik · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Later among the works it cites.
Active image segmentation propagation
Suyog Dutt Jain and Kristen Grauman · 2016
Later among the works it cites.
Learning transferrable knowledge for semantic segmentation with deep convolutional neural network
Seunghoon Hong, Junhyuk Oh, Honglak Lee, and Bohyung Han · 2016
Later among the works it cites.
Text-guided attention model for image captioning
Jonghwan Mun, Minsu Cho, and Bohyung Han · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Enriching visual knowledge bases via object discovery and segmentation
Xinlei Chen, Abhinav Shrivastava, and Abhinav Gupta · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Cited alongside, same era.
Recurrent models of visual attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, et al · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Look and think twice: Capturing top-down visual attention with feedback convolutional neural networks
Chunshui Cao, Xianming Liu, Yi Yang, Yinan Yu, Jiang Wang, Zilei Wang, Yongzhen Huang, Liang Wang, Chang Huang, Wei Xu, Deva Ramanan, and Thomas S. Huang · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al · 2015
Cited alongside, same era.
Hierarchical attention networks
Paul Hongsuck Seo, Zhe Lin, Scott Cohen, Xiaohui Shen, and Bohyung Han · 2016
Later among the works it cites.
A theoretical framework for robustness of (deep) classifiers under adversarial noise
Beilun Wang, Ji Gao, and Yanjun Qi · 2016
Later among the works it cites.
Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering
Huijuan Xu and Kate Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering, 06 2016
Zichao Yang, Xiaodong He, Jianfeng Gao, li Deng, and Alex Smola · 2016
Later among the works it cites.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Later among the works it cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Later among the works it cites.
Deepmask: Masking DNN models for robustness against adversarial samples
Ji Gao, Beilun Wang, and Yanjun Qi · 2017
Later among the works it cites.