Fetching the paper…
Reading the bibliography…
In this paper we approach the novel problem of segmenting an image based on a natural language expression.
Long short-term memory
Hochreiter, S., Schmidhuber, J.: · 1997
Earlier work this paper cites.
Grabcut: Interactive foreground extraction using iterated graph cuts
Rother, C., Kolmogorov, V., Blake, A.: · 2004
Earlier work this paper cites.
The iapr tc-12 benchmark: A new evaluation resource for visual information systems
Grubinger, M., Clough, P., Müller, H., Deselaers, T.: · 2006
Earlier work this paper cites.
The segmented and annotated iapr tc-12 benchmark
Escalante, H.J., Hernández, C.A., Gonzalez, J.A., López-López, A., Montes, M., Morales, E.F., Sucar, L.E., Villaseñor, L., Grubinger, M.: · 2010
Earlier work this paper cites.
Semantic segmentation with second-order pooling
Carreira, J., Caseiro, R., Batista, J., Sminchisescu, C.: · 2012
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A.: · 2012
Earlier work this paper cites.
Semantic image segmentation with deep convolutional nets and fully connected crfs
Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: · 2014
Earlier work this paper cites.
Simultaneous detection and segmentation
Hariharan, B., Arbeláez, P., Girshick, R., Malik, J.: · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., Le, Q.V.: · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Cho, K., van Merriënboer, B., Bahdanau, D., Bengio, Y.: · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., Berg, T.L.: · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: · 2014
Cited alongside, same era.
Edge boxes: Locating object proposals from edges
Zitnick, C.L., Dollár, P.: · 2014
Cited alongside, same era.
Multiscale combinatorial grouping
Arbeláez, P., Pont-Tuset, J., Barron, J., Marques, F., Malik, J.: · 2014
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., Darrell, T.: · 2015
Cited alongside, same era.
Conditional random fields as recurrent neural networks
Zheng, S., Jayasumana, S., Romera-Paredes, B., Vineet, V., Su, Z., Du, D., Huang, C., Torr, P.H.: · 2015
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
Rohrbach, A., Rohrbach, M., Hu, R., Darrell, T., Schiele, B.: · 2015
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A., Murphy, K.: · 2015
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., Darrell, T.: · 2015
Later among the works it cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Mao, J., Xu, W., Yang, Y., Wang, J., Huang, Z., Yuille, A.: · 2015
Later among the works it cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B., Wang, L., Cervantes, C., Caicedo, J., Hockenmaier, J., Lazebnik, S.: · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning deconvolution network for semantic segmentation
Noh, H., Hong, S., Han, B.: · 2015
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
Yu, F., Koltun, V.: · 2015
Cited alongside, same era.
Im2calories: towards an automated mobile vision food diary
Meyers, A., Johnston, N., Rathod, V., Korattikara, A., Gorban, A., Silberman, N., Guadarrama, S., Papandreou, G., Huang, J., Murphy, K.P.: · 2015
Cited alongside, same era.
Natural language object retrieval
Hu, R., Xu, H., Rohrbach, M., Feng, J., Saenko, K., Darrell, T.: · 2015
Cited alongside, same era.
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., Bengio, Y.: · 2015
Later among the works it cites.
Stacked attention networks for image question answering
Yang, Z., He, X., Gao, J., Deng, L., Smola, A.: · 2015
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Xu, H., Saenko, K.: · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al.: · 2015
Later among the works it cites.
Learning to compose neural networks for question answering
Andreas, J., Rohrbach, M., Darrell, T., Klein, D.: · 2016
Closest in time.