Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel end-to-end model, namely Single-Stage Grounding network (SSG), to localize the referent given a referring expression within an image.
Long short-term memory
S. Hochreiter and Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
The segmented and annotated iapr tc-12 benchmark
H. J. Escalante, C. A. Hernández, J. A. Gonzalez, A. López-López, M. Montes, E. F. Morales, L. Enrique Sucar, L. Villaseñor, and M. Grubinger · 2010
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas, A. Y. Hannun, and A. Y. Ng · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Selective search for object recognition
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Referit game: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Reconciling saliency and object center-bias hypotheses in explaining free-viewing fixations
A. Borji and J. Tanner · 2016
Cited alongside, same era.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Cited alongside, same era.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Cited alongside, same era.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, and K. Murphy · 2016
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Later among the works it cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Later among the works it cites.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Later among the works it cites.
Referring expression generation and comprehension via attributes
J. Liu, L. Wang, and M.-H. Yang · 2017
Later among the works it cites.
Video captioning with transferred semantic attributes
Y. Pan, T. Yao, H. Li, and T. Mei · 2017
Later among the works it cites.
Boosting image captioning with attributes
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Cited alongside, same era.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Cited alongside, same era.
Yolo9000: Better, faster, stronger
J. Redmon and A. Farhadi · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi · 2016
Cited alongside, same era.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Cited alongside, same era.
A joint speaker-listener-reinforcer model for referring expressions
L. Yu, H. Tan, M. Bansal, and T. L. Berg · 2017
Later among the works it cites.
Learning to guide decoding for image captioning
W. Jiang, L. Ma, X. Chen, H. Zhang, and W. Liu · 2018
Closest in time.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Closest in time.
Yolov3: An incremental improvement
J. Redmon and A. Farhadi · 2018
Closest in time.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Closest in time.
Grounding referring expressions in images by variational context
H. Zhang, Y. Niu, and S.-F. Chang · 2018
Closest in time.