Fetching the paper…
Reading the bibliography…
Different from universal object detection, referring expression comprehension (REC) aims to locate specific objects referred to by natural language expressions.
Batch-shaping for learning conditional channel gated networks
Bejnordi, B. E.; Blankevoort, T.; and Welling, M. 2019 · 1907
Earlier work this paper cites.
The segmented and annotated IAPR TC-12 benchmark
Escalante, H. J.; Hernández, C. A.; Gonzalez, J. A.; López-López, A.; Montes, M.; Morales, E. F.; Sucar, L. E.; Villasenor, L.; and Grubinger, M. 2010 · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X.; and Bengio, Y. 2010 · 2010
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E.; Gu, S.; and Poole, B. 2016 · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J.; Mnih, A.; and Teh, Y. W. 2016 · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J.; Huang, J.; Toshev, A.; Camburu, O.; Yuille, A. L.; and Murphy, K. 2016 · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Nagaraja, V. K.; Morariu, V. I.; and Davis, L. S. 2016 · 2016
Earlier work this paper cites.
Faster r-cnn features for instance search
Salvador, A.; Giró-i Nieto, X.; Marqués, F.; and Satoh, S. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Referring expression generation and comprehension via attributes
Liu, J.; Wang, L.; and Yang, M.-H. 2017 · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; and Dean, J. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Dynamic channel pruning: Feature boosting and suppression
Gao, X.; Zhao, Y.; Dudziak, Ł.; Mullins, R.; and Xu, C.-z. 2018 · 2018
Cited alongside, same era.
A fast and accurate one-stage approach to visual grounding
Yang, Z.; Gong, B.; Wang, L.; Huang, W.; Yu, D.; and Luo, J. 2019 · 2019
Later among the works it cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Later among the works it cites.
Uniter: Universal image-text representation learning
Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Later among the works it cites.
Large-scale adversarial training for vision-and-language representation learning
Gan, Z.; Chen, Y.-C.; Li, L.; Zhu, C.; Cheng, Y.; and Liu, J. 2020 · 2020
Later among the works it cites.
Channel selection using gumbel softmax
Herrmann, C.; Bowen, R. S.; and Zabih, R. 2020 · 2020
Later among the works it cites.
A real-time cross-modality correlation filtering method for referring expression comprehension
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaiser, Ł.; and Bengio, S. 2018 · 2018
Cited alongside, same era.
Tell-and-answer: Towards explainable visual question answering using attributes and captions
Li, Q.; Fu, J.; Yu, D.; Mei, T.; and Luo, J. 2018 · 2018
Cited alongside, same era.
Convolutional networks with adaptive inference graphs
Veit, A.; and Belongie, S. 2018 · 2018
Cited alongside, same era.
Skipnet: Learning dynamic routing in convolutional networks
Wang, X.; Yu, F.; Dou, Z.-Y.; Darrell, T.; and Gonzalez, J. E. 2018 · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Yu, L.; Lin, Z.; Shen, X.; Yang, J.; Lu, X.; Bansal, M.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
G3raphground: Graph-based language grounding
Bajaj, M.; Wang, L.; and Sigal, L. 2019 · 2019
Cited alongside, same era.
Once-for-All: Train One Network and Specialize it for Efficient Deployment
Cai, H.; Gan, C.; Wang, T.; Zhang, Z.; and Han, S. 2019 · 2019
Cited alongside, same era.
Learning to compose and reason with language tree structures for visual grounding
Hong, R.; Liu, D.; Mo, X.; He, X.; and Zhang, H. 2019 · 2019
Cited alongside, same era.
Liao, Y.; Liu, S.; Li, G.; Wang, F.; Chen, Y.; Qian, C.; and Li, B. 2020 · 2020
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2020 · 2020
Later among the works it cites.
Multi-task collaborative network for joint referring expression comprehension and segmentation
Luo, G.; Zhou, Y.; Sun, X.; Cao, L.; Wu, C.; Deng, C.; and Ji, R. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2020 · 2020
Later among the works it cites.
Improving one-stage visual grounding by recursive sub-query construction
Yang, Z.; Chen, T.; Wang, L.; and Luo, J. 2020 · 2020
Later among the works it cites.
Ref-NMS: breaking proposal bottlenecks in two-stage referring expression grounding
Chen, L.; Ma, W.; Xiao, J.; Zhang, H.; and Chang, S.-F. 2021 · 2021
Later among the works it cites.
Transvg: End-to-end visual grounding with transformers
Deng, J.; Yang, Z.; Chen, T.; Zhou, W.; and Li, H. 2021 · 2021
Later among the works it cites.
MDETR-modulated detection for end-to-end multi-modal understanding
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Later among the works it cites.
Fully dynamic inference with deep neural networks
Xia, W.; Yin, H.; Dai, X.; and Jha, N. K. 2021 · 2021
Later among the works it cites.
A real-time global inference network for one-stage referring expression comprehension
Zhou, Y.; Ji, R.; Luo, G.; Sun, X.; Su, J.; Ding, X.; Lin, C.-W.; and Tian, Q. 2021 · 2021
Later among the works it cites.