Fetching the paper…
Reading the bibliography…
Panoptic Narrative Detection (PND) and Segmentation (PNS) are two challenging tasks that involve identifying and locating multiple targets in an image according to a long narrative description.
Liu X, Wang Z, Shao J, et al (2019b) Improving referring expression grounding with cross-modal attention-guided erasing. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 1950–1959
1959
Earlier work this paper cites.
Girshick R, Donahue J, Darrell T, et al (2014) Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 580–587
2014
Earlier work this paper cites.
Karpathy A, Joulin A, Fei-Fei LF (2014) Deep fragment embeddings for bidirectional image sentence mapping. Advances in neural information processing systems 27
2014
Earlier work this paper cites.
Kazemzadeh S, Ordonez V, Matten M, et al (2014) Referitgame: Referring to objects in photographs of natural scenes. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp 787–798
2014
Earlier work this paper cites.
Lin TY, Maire M, Belongie S, et al (2014) Microsoft coco: Common objects in context. In: European conference on computer vision, Springer, pp 740–755
2014
Earlier work this paper cites.
Girshick R (2015) Fast r-cnn. In: Proceedings of the IEEE international conference on computer vision, pp 1440–1448
2015
Earlier work this paper cites.
Plummer BA, Wang L, Cervantes CM, et al (2015) Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In: Proceedings of the IEEE international conference on computer vision, pp 2641–2649
2015
Earlier work this paper cites.
Ba JL, Kiros JR, Hinton GE (2016) Layer normalization. arXiv preprint arXiv:160706450
2016
Earlier work this paper cites.
He K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770–778
2016
Earlier work this paper cites.
Mao J, Huang J, Toshev A, et al (2016) Generation and comprehension of unambiguous object descriptions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 11–20
2016
Earlier work this paper cites.
Milletari F, Navab N, Ahmadi SA (2016) V-net: Fully convolutional neural networks for volumetric medical image segmentation. In: 2016 fourth international conference on 3D vision (3DV), IEEE, pp 565–571
2016
Earlier work this paper cites.
Nagaraja VK, Morariu VI, Davis LS (2016) Modeling context between objects for referring expression understanding. In: European Conference on Computer Vision, Springer, pp 792–807
2016
Earlier work this paper cites.
Redmon J, Divvala S, Girshick R, et al (2016) You only look once: Unified, real-time object detection. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 779–788
2016
Earlier work this paper cites.
Yang Z, Yang D, Dyer C, et al (2016) Hierarchical attention networks for document classification. In: Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies, pp 1480–1489
2016
Earlier work this paper cites.
Yu L, Poirson P, Yang S, et al (2016) Modeling context in referring expressions. In: European Conference on Computer Vision, Springer, pp 69–85
2016
Earlier work this paper cites.
He K, Gkioxari G, Dollár P, et al (2017) Mask r-cnn. In: Proceedings of the IEEE international conference on computer vision, pp 2961–2969
2017
Earlier work this paper cites.
Hu R, Rohrbach M, Andreas J, et al (2017) Modeling relationships in referential expressions with compositional modular networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1115–1124
2017
Earlier work this paper cites.
Lin TY, Dollár P, Girshick R, et al (2017) Feature pyramid networks for object detection. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2117–2125
2017
Earlier work this paper cites.
Luo R, Shakhnarovich G (2017) Comprehension-guided referring expressions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 7102–7111
2017
Earlier work this paper cites.
Redmon J, Farhadi A (2017) Yolo9000: better, faster, stronger. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 7263–7271
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, et al (2017) Attention is all you need. Advances in neural information processing systems 30
2017
Cited alongside, same era.
Jiang B, Luo R, Mao J, et al (2018) Acquisition of localization confidence for accurate object detection. In: Proceedings of the European conference on computer vision (ECCV), pp 784–799
2018
Cited alongside, same era.
Li R, Li K, Kuo YC, et al (2018) Referring image segmentation via recurrent refinement networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 5745–5753
2018
Cited alongside, same era.
Margffoy-Tuay E, Pérez JC, Botero E, et al (2018) Dynamic multimodal instance segmentation guided by natural language queries. In: Proceedings of the European Conference on Computer Vision (ECCV), pp 630–645
2018
Cited alongside, same era.
Jing C, Wu Y, Pei M, et al (2020) Visual-semantic graph matching for visual grounding. In: Proceedings of the 28th ACM International Conference on Multimedia, pp 4041–4050
2020
Later among the works it cites.
Luo G, Zhou Y, Sun X, et al (2020) Multi-task collaborative network for joint referring expression comprehension and segmentation. In: Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pp 10034–10043
2020
Later among the works it cites.
Pont-Tuset J, Uijlings J, Changpinyo S, et al (2020) Connecting vision and language with localized narratives. In: European conference on computer vision, Springer, pp 647–664
2020
Later among the works it cites.
Yu T, Hui T, Yu Z, et al (2020) Cross-modal omni interaction modeling for phrase grounding. In: Proceedings of the 28th ACM International Conference on Multimedia, pp 1725–1734
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Plummer BA, Kordas P, Kiapour MH, et al (2018) Conditional image-text embedding networks. In: Proceedings of the European Conference on Computer Vision (ECCV), pp 249–264
2018
Cited alongside, same era.
Redmon J, Farhadi A (2018) Yolov3: An incremental improvement. arXiv preprint arXiv:180402767
2018
Cited alongside, same era.
Shi H, Li H, Meng F, et al (2018) Key-word-aware network for referring expression image segmentation. In: Proceedings of the European Conference on Computer Vision (ECCV), pp 38–54
2018
Cited alongside, same era.
Wang L, Li Y, Huang J, et al (2018) Learning two-branch neural networks for image-text matching tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2):394–407
2018
Cited alongside, same era.
Yu L, Lin Z, Shen X, et al (2018) Mattnet: Modular attention network for referring expression comprehension. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 1307–1315
2018
Cited alongside, same era.
Akbari H, Karaman S, Bhargava S, et al (2019) Multi-level multimodal common semantic space for image-phrase grounding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 12476–12486
2019
Cited alongside, same era.
Bajaj M, Wang L, Sigal L (2019) G3raphground: Graph-based language grounding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 4281–4290
2019
Cited alongside, same era.
Kenton JDMWC, Toutanova LK (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT, pp 4171–4186
2019
Cited alongside, same era.
Ding H, Liu C, Wang S, et al (2021) Vision-language transformer and query generation for referring segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 16321–16330
2021
Later among the works it cites.
González C, Ayobi N, Hernández I, et al (2021) Panoptic narrative grounding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 1364–1373
2021
Later among the works it cites.
Kamath A, Singh M, LeCun Y, et al (2021) Mdetr-modulated detection for end-to-end multi-modal understanding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 1780–1790
2021
Later among the works it cites.
Li M, Sigal L (2021) Referring transformer: A one-step approach to multi-task visual grounding. Advances in Neural Information Processing Systems 34:19652–19664
2021
Later among the works it cites.
Liu S, Hui T, Huang S, et al (2021) Cross-modal progressive comprehension for referring segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 44(9):4761–4775
2021
Later among the works it cites.
Mu Z, Tang S, Tan J, et al (2021) Disentangled motif-aware graph learning for phrase grounding. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp 13587–13594
2021
Later among the works it cites.
Zhang W, Pang J, Chen K, et al (2021) K-net: Towards unified image segmentation. Advances in Neural Information Processing Systems 34:10326–10338
2021
Later among the works it cites.
Zhou Y, Ji R, Luo G, et al (2021) A real-time global inference network for one-stage referring expression comprehension. IEEE Transactions on Neural Networks and Learning Systems
2021
Later among the works it cites.
Chen YW, Tsai YH, Yang MH (2022) Understanding synonymous referring expressions via contrastive features. International Journal of Computer Vision 130(10):2501–2516
2022
Later among the works it cites.
Cheng B, Misra I, Schwing AG, et al (2022) Masked-attention mask transformer for universal image segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 1290–1299
2022
Later among the works it cites.
Ding Z, Ding Zh, Hui T, et al (2022) Ppmn: Pixel-phrase matching network for one-stage panoptic narrative grounding. In: Proceedings of the 30th ACM International Conference on Multimedia, pp 5537–5546
2022
Later among the works it cites.
González C, Ayobi N, Hernández I, et al (2023) Piglet: Pixel-level grounding of language expressions with transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence
2023
Closest in time.
Hui T, Ding Z, Huang J, et al (2023) Enriching phrases with coupled pixel and object contexts for panoptic narrative grounding. In: Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI 2023, 19th-25th August 2023, Macao, SAR, China, pp 893–901
2023
Closest in time.
Luo G, Zhou Y, Sun X, et al (2023) Towards language-guided visual recognition via dynamic convolutions. International Journal of Computer Vision pp 1–19
2023
Closest in time.
Wang H, Ji J, Zhou Y, et al (2023) Towards real-time panoptic narrative grounding by an end-to-end grounding network. Proceedings of the AAAI Conference on Artificial Intelligence
2023
Closest in time.