Fetching the paper…
Reading the bibliography…
Gaze reflects how humans process visual scenes and is therefore increasingly used in computer vision systems.
A. L. Yarbus, Eye movements and vision . Springer, 1967
1967
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proc. ACL , 2002, pp. 311–318
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out: Proceedings of the ACL-04 workshop , vol. 8, 2004
2004
Earlier work this paper cites.
W. Einhäuser, M. Spain, and P. Perona, “Objects predict fixations better than early saliency,” Journal of Vision , vol. 8, no. 14, p. 18, 2008
2008
Earlier work this paper cites.
S. Bird, E. Klein, and E. Loper, Natural language processing with Python . O’Reilly Media Inc., 2009
2009
Earlier work this paper cites.
H. Larochelle and G. E. Hinton, “Learning to combine foveal glimpses with a third-order boltzmann machine,” in Proc. NIPS , 2010, pp. 1243–1251
2010
Earlier work this paper cites.
A. Nuthmann and J. M. Henderson, “Object-based attentional selection in scene viewing,” Journal of vision , vol. 10, no. 8, p. 20, 2010
2010
Earlier work this paper cites.
S. Ramanathan, H. Katti, N. Sebe, M. Kankanhalli, and T.-S. Chua, “An eye fixation database for saliency detection in images,” in Proc. ECCV , 2010, pp. 30–43
2010
Earlier work this paper cites.
A. Bulling, J. A. Ward, H. Gellersen, and G. Tröster, “Eye movement analysis for activity recognition using electrooculography,” IEEE TPAMI , vol. 33, no. 4, pp. 741–753, 2011
2011
Earlier work this paper cites.
S. Ramanathan, V. Yanulevskaya, and N. Sebe, “Can computers learn from humans to see better?: inferring scene semantics from viewers’ eye movements,” in Proc. ACMMM , 2011, pp. 33–42
2011
Earlier work this paper cites.
A. Fathi, Y. Li, and J. M. Rehg, “Learning to recognize daily actions using gaze,” in Proc. ECCV , 2012, pp. 314–327
2012
Earlier work this paper cites.
D. Rudoy, D. B. Goldman, E. Shechtman, and L. Zelnik-Manor, “Crowdsourcing gaze data collection,” in Proc. Collective Intelligence , 2012
2012
Earlier work this paper cites.
A. K. Mishra, Y. Aloimonos, L. F. Cheong, and A. Kassim, “Active visual segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 34, no. 4, pp. 639–653, April 2012
2012
Earlier work this paper cites.
T. Toyama, T. Kieninger, F. Shafait, and A. Dengel, “Gaze guided object recognition using a head-mounted eye tracker,” in Proc. ETRA , 2012, pp. 91–98
2012
Earlier work this paper cites.
——, “Scan patterns predict sentence production in the cross-modal processing of visual scenes,” Cognitive Science , vol. 36, no. 7, pp. 1204–1223, 2012
2012
Earlier work this paper cites.
S. Karthikeyan, V. Jagadeesh, R. Shenoy, M. Eckstein, and B. Manjunath, “From where and how to what we see,” in Proc. ICCV , 2013, pp. 625–632
2013
Earlier work this paper cites.
K. Yun, Y. Peng, D. Samaras, G. J. Zelinsky, and T. L. Berg, “Exploring the role of gaze behavior and object detection in scene understanding,” Frontiers in psychology , vol. 4, 2013
2013
Earlier work this paper cites.
G. J. Zelinsky, “Understanding scene understanding,” Frontiers in psychology , vol. 4, 2013
2013
Earlier work this paper cites.
A. Borji and L. Itti, “State-of-the-art in visual attention modeling,” IEEE TPAMI , vol. 35, no. 1, pp. 185–207, 2013
2013
Earlier work this paper cites.
Y. Sugano, Y. Matsushita, and Y. Sato, “Graph-based joint clustering of fixations and visual entities,” ACM Transactions on Applied Perception (TAP) , vol. 10, no. 2, p. 10, 2013
2013
Earlier work this paper cites.
K. Yun, Y. Peng, D. Samaras, G. J. Zelinsky, and T. Berg, “Studying relationships between human gaze, description, and computer vision,” in Proc. CVPR . IEEE, 2013, pp. 739–746
2013
Earlier work this paper cites.
D. P. Papadopoulos, A. D. Clarke, F. Keller, and V. Ferrari, “Training object class detectors from eye tracking data,” in Proc. ECCV , 2014, pp. 361–376
2014
Earlier work this paper cites.
K. A. Funes Mora and J.-M. Odobez, “Geometric generative gaze estimation (G3E) for remote RGB-D cameras,” in Proc. CVPR , 2014, pp. 1773–1780
2014
Cited alongside, same era.
Y. Sugano, Y. Matsushita, and Y. Sato, “Learning-by-synthesis for appearance-based 3d gaze estimation,” in Proc. CVPR . IEEE, 2014, pp. 1821–1828
2014
Cited alongside, same era.
V. Mnih, N. Heess, A. Graves et al. , “Recurrent models of visual attention,” in Proc. NIPS , 2014, pp. 2204–2212
2014
Cited alongside, same era.
Y. Zheng, R. S. Zemel, Y.-J. Zhang, and H. Larochelle, “A neural autoregressive approach to attention-based recognition,” IJCV , vol. 113, no. 1, pp. 67–79, 2014
2014
Cited alongside, same era.
J. L. Long, N. Zhang, and T. Darrell, “Do convnets learn correspondence?” in Proc. NIPS , 2014, pp. 1601–1609
2014
Cited alongside, same era.
S. Mathe and C. Sminchisescu, “Actions in the eye: dynamic gaze datasets and learnt saliency models for visual recognition,” IEEE TPAMI , vol. 37, no. 7, pp. 1408–1424, 2015
2015
Later among the works it cites.
M. I. Coco and F. Keller, “Integrating mechanisms of visual guidance in naturalistic language production,” Cognitive processing , vol. 16, no. 2, pp. 131–150, 2015
2015
Later among the works it cites.
J. Ba, V. Mnih, and K. Kavukcuoglu, “Multiple object recognition with visual attention,” in Proc. ICLR , 2015
2015
Later among the works it cites.
K. Gregor, I. Danihelka, A. Graves, and D. Wierstra, “DRAW: A recurrent neural network for image generation,” in Proc. ICML , 2015
2015
Later among the works it cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. ICLR , 2015
2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, “The secrets of salient object segmentation,” in Proc. CVPR , 2014, pp. 280–287
2014
Cited alongside, same era.
J. Xu, M. Jiang, S. Wang, M. S. Kankanhalli, and Q. Zhao, “Predicting human gaze beyond pixels,” Journal of vision , vol. 14, no. 1, p. 28, 2014
2014
Cited alongside, same era.
M. Ranzato, “On learning where to look,” arXiv preprint arXiv:1405.5488 , 2014
2014
Cited alongside, same era.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” IJCV , pp. 1–42, 2014
2014
Cited alongside, same era.
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in Proc. NIPS , 2014, pp. 487–495
2014
Cited alongside, same era.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in Proc. ECCV , 2014, pp. 740–755
2014
Cited alongside, same era.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in Proc. ECCV , 2014, pp. 818–833
2014
Cited alongside, same era.
Later among the works it cites.
2015
Later among the works it cites.
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell, “Language models for image captioning: The quirks and what works,” in Proc. ACL , 2015
2015
Later among the works it cites.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proc. CVPR , 2015
2015
Later among the works it cites.
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Doll¥’ar, J. Gao, X. He, M. Mitchell, J. Platt et al. , “From captions to visual concepts and back,” in Proc. CVPR , 2015
2015
Later among the works it cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proc. CVPR , 2015
2015
Later among the works it cites.
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille, “Learning like a child: Fast novel visual concept learning from sentence descriptions of images,” in Proc. ICCV , 2015
2015
Later among the works it cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proc. CVPR , 2015
2015
Later among the works it cites.
X. Chen and C. Lawrence Zitnick, “Mind’s eye: A recurrent visual representation for image caption generation,” in Proc. CVPR , 2015, pp. 2422–2431
2015
Later among the works it cites.
Y. Ushiku, M. Yamaguchi, Y. Mukuta, and T. Harada, “Common subspace for model and similarity: Phrase learning for caption generation from images,” in Proc. ICCV , 2015, pp. 2668–2676
2015
Later among the works it cites.
J. Zhang and S. Sclaroff, “Exploiting surroundedness for saliency detection: a boolean map approach,” IEEE TPAMI , 2015
2015
Later among the works it cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. ICLR , 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
Y.-W. Chao, Z. Wang, R. Mihalcea, and J. Deng, “Mining semantic affordances of visual object categories,” in Proc. CVPR , 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
N. Liu, J. Han, D. Zhang, S. Wen, and T. Liu, “Predicting eye fixations using convolutional neural networks,” in Proc. CVPR , 2015, pp. 362–370
2015
Later among the works it cites.
2015
Later among the works it cites.
D. Damen, T. Leelasawassuk, and W. Mayol-Cuevas, “You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance,” Computer Vision and Image Understanding , vol. 149, pp. 98 – 112, 2016
2016
Closest in time.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo, “Image captioning with semantic attention,” in Proc. CVPR , 2016
2016
Closest in time.