Fetching the paper…
Reading the bibliography…
Panoptic Narrative Grounding (PNG) is an emerging cross-modal grounding task, which locates the target regions of an image corresponding to the text description.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
Hu, R.; Rohrbach, M.; and Darrell, T. 2016 · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J.; Huang, J.; Toshev, A.; Camburu, O.; Yuille, A. L.; and Murphy, K. 2016 · 2016
Earlier work this paper cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Milletari, F.; Navab, N.; and Ahmadi, S.-A. 2016 · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Nagaraja, V. K.; Morariu, V. I.; and Davis, L. S. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017 · 2017
Earlier work this paper cites.
Recurrent multimodal interaction for referring image segmentation
Liu, C.; Lin, Z.; Shen, X.; Yang, J.; Lu, X.; and Yuille, A. 2017 · 2017
Earlier work this paper cites.
Referring expression generation and comprehension via attributes
Liu, J.; Wang, L.; and Yang, M.-H. 2017 · 2017
Earlier work this paper cites.
Comprehension-guided referring expressions
Luo, R.; and Shakhnarovich, G. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
A joint speaker-listener-reinforcer model for referring expressions
Yu, L.; Tan, H.; Bansal, M.; and Berg, T. L. 2017 · 2017
Cited alongside, same era.
Panoptic Segmentation with a Joint Semantic and Instance Segmentation Network
de Geus, D.; Meletis, P.; and Dubbelman, G. 2018 · 2018
Cited alongside, same era.
Referring image segmentation via recurrent refinement networks
Li, R.; Li, K.; Kuo, Y.-C.; Shu, M.; Qi, X.; Shen, X.; and Jia, J. 2018 · 2018
Cited alongside, same era.
Dynamic multimodal instance segmentation guided by natural language queries
Margffoy-Tuay, E.; Pérez, J. C.; Botero, E.; and Arbeláez, P. 2018 · 2018
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2020
Later among the works it cites.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020 · 2020
Later among the works it cites.
Per-pixel classification is not all you need for semantic segmentation
Cheng, B.; Schwing, A.; and Kirillov, A. 2021 · 2021
Later among the works it cites.
Panoptic Narrative Grounding
González, C.; Ayobi, N.; Hernández, I.; Hernández, J.; Pont-Tuset, J.; and Arbeláez, P. 2021 · 2021
Later among the works it cites.
Improving image captioning by leveraging intra-and inter-layer global representation in transformer network
Ji, J.; Luo, Y.; Sun, X.; Chen, F.; Luo, G.; Wu, Y.; Gao, Y.; and Ji, R. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shi, H.; Li, H.; Meng, F.; and Wu, Q. 2018 · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Yu, L.; Lin, Z.; Shen, X.; Yang, J.; Lu, X.; Bansal, M.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Kenton, J. D. M.-W. C.; and Toutanova, L. K. 2019 · 2019
Cited alongside, same era.
Learning to assemble neural module tree networks for visual grounding
Liu, D.; Zhang, H.; Wu, F.; and Zha, Z.-J. 2019 · 2019
Cited alongside, same era.
Upsnet: A unified panoptic segmentation network
Xiong, Y.; Liao, R.; Zhao, H.; Hu, R.; Bai, M.; Yumer, E.; and Urtasun, R. 2019 · 2019
Cited alongside, same era.
Cross-modal self-attention network for referring image segmentation
Ye, L.; Rochan, M.; Liu, Z.; and Wang, Y. 2019 · 2019
Cited alongside, same era.
Deep modular co-attention networks for visual question answering
Yu, Z.; Yu, J.; Cui, Y.; Tao, D.; and Tian, Q. 2019 · 2019
Cited alongside, same era.
MDETR-modulated detection for end-to-end multi-modal understanding
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Later among the works it cites.
Fully convolutional networks for panoptic segmentation
Li, Y.; Zhao, H.; Qi, X.; Wang, L.; Li, Z.; Sun, J.; and Jia, J. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Max-deeplab: End-to-end panoptic segmentation with mask transformers
Wang, H.; Zhu, Y.; Adam, H.; Yuille, A.; and Chen, L.-C. 2021 · 2021
Later among the works it cites.
K-net: Towards unified image segmentation
Zhang, W.; Pang, J.; Chen, K.; and Loy, C. C. 2021 · 2021
Later among the works it cites.
A real-time global inference network for one-stage referring expression comprehension
Zhou, Y.; Ji, R.; Luo, G.; Sun, X.; Su, J.; Ding, X.; Lin, C.-W.; and Tian, Q. 2021 · 2021
Later among the works it cites.
Panoptic SegFormer: Delving deeper into panoptic segmentation with transformers
Li, Z.; Wang, W.; Xie, E.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; Luo, P.; and Lu, T. 2022 · 2022
Later among the works it cites.
DA-Transformer: Distance-aware Transformer
Wu, C.; Wu, F.; and Huang, Y. 2021 · 2068
Closest in time.