Fetching the paper…
Reading the bibliography…
Unified visual grounding pursues a simple and generic technical route to leverage multi-task data with less task-specific design.
A characterization of ten hidden-surface algorithms
Ivan E Sutherland, Robert F Sproull, and Robert A Schumacker · 1974
Earlier work this paper cites.
Topological structural analysis of digitized binary images by border following
Satoshi Suzuki et al · 1985
Earlier work this paper cites.
The opencv library
Gary Bradski · 2000
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
Ronghang Hu, Marcus Rohrbach, and Trevor Darrell · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Modeling relationships in referential expressions with compositional modular networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Referring expression generation and comprehension via attributes
Jingyu Liu, Liang Wang, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Comprehension-guided referring expressions
Ruotian Luo and Gregory Shakhnarovich · 2017
Earlier work this paper cites.
A joint speaker-listener-reinforcer model for referring expressions
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg · 2017
Earlier work this paper cites.
Discriminative bimodal networks for visual localization and detection with natural language queries
Yuting Zhang, Luyao Yuan, Yijie Guo, Zhiyuan He, I-An Huang, and Honglak Lee · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Cornernet: Detecting objects as paired keypoints
Hei Law and Jia Deng · 2018
Earlier work this paper cites.
Referring image segmentation via recurrent refinement networks
Ruiyu Li, Kaican Li, Yi-Chun Kuo, Michelle Shu, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia · 2018
Earlier work this paper cites.
Dynamic multimodal instance segmentation guided by natural language queries
Edgar Margffoy-Tuay, Juan C Pérez, Emilio Botero, and Pablo Arbeláez · 2018
Earlier work this paper cites.
Yolov3: An incremental improvement
Joseph Redmon and Ali Farhadi · 2018
Earlier work this paper cites.
Key-word-aware network for referring expression image segmentation
Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Grounding referring expressions in images by variational context
Hanwang Zhang, Yulei Niu, and Shih-Fu Chang · 2018
Earlier work this paper cites.
Parallel attention: A unified framework for visual object discovery through dialogs and queries
Bohan Zhuang, Qi Wu, Chunhua Shen, Ian Reid, and Anton Van Den Hengel · 2018
Earlier work this paper cites.
See-through-text grouping for referring image segmentation
Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, and Tyng-Luh Liu · 2019
Earlier work this paper cites.
Referring expression object segmentation with caption-aware consistency
Yi-Wen Chen, Yi-Hsuan Tsai, Tiantian Wang, Yen-Yu Lin, and Ming-Hsuan Yang · 2019
Earlier work this paper cites.
Learning to compose and reason with language tree structures for visual grounding
Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He, and Hanwang Zhang · 2019
Cited alongside, same era.
Learning to assemble neural module tree networks for visual grounding
Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha · 2019
Cited alongside, same era.
Improving referring expression grounding with cross-modal attention-guided erasing
Xihui Liu, Zihao Wang, Jing Shao, Xiaogang Wang, and Hongsheng Li · 2019
Cited alongside, same era.
Fcos: Fully convolutional one-stage object detection
Zhi Tian, Chunhua Shen, Hao Chen, and Tong He · 2019
Cited alongside, same era.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel · 2019
Cited alongside, same era.
Dynamic graph attention for referring expression comprehension
Transvg: End-to-end visual grounding with transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li · 2021
Later among the works it cites.
Vision-language transformer and query generation for referring segmentation
Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang · 2021
Later among the works it cites.
Encoder fusion network with co-attention embedding for referring image segmentation
Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu · 2021
Later among the works it cites.
Look before you leap: Learning landmark features for one-stage visual grounding
Binbin Huang, Dongze Lian, Weixin Luo, and Shenghua Gao · 2021
Later among the works it cites.
Referring transformer: A one-step approach to multi-task visual grounding
Muchen Li and Leonid Sigal · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sibei Yang, Guanbin Li, and Yizhou Yu · 2019
Cited alongside, same era.
A fast and accurate one-stage approach to visual grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo · 2019
Cited alongside, same era.
Cross-modal self-attention network for referring image segmentation
Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang · 2019
Cited alongside, same era.
Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Learning to branch for multi-task learning
Pengsheng Guo, Chen-Yu Lee, and Daniel Ulbricht · 2020
Cited alongside, same era.
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Conditional detr for fast training convergence
Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang · 2021
Later among the works it cites.
Iterative shrinking for referring expression grounding using deep reinforcement learning
Mingjie Sun, Jimin Xiao, and Eng Gee Lim · 2021
Later among the works it cites.
Bottom-up shift and reasoning for referring image segmentation
Sibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou, and Yizhou Yu · 2021
Later among the works it cites.
Lavt: Language-aware vision transformer for referring image segmentation
Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr · 2021
Later among the works it cites.
Trar: Routing the attention spans in transformer for visual question answering
Yiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun, Jianzhuang Liu, Xinghao Ding, Mingliang Xu, and Rongrong Ji · 2021
Later among the works it cites.
A generalist framework for panoptic segmentation of images and videos
Ting Chen, Lala Li, Saurabh Saxena, Geoffrey Hinton, and David J Fleet · 2022
Later among the works it cites.
A unified sequence interface for vision tasks
Ting Chen, Saurabh Saxena, Lala Li, Tsung-Yi Lin, David J Fleet, and Geoffrey Hinton · 2022
Later among the works it cites.
Analog bits: Generating discrete data using diffusion models with self-conditioning
Ting Chen, Ruixiang Zhang, and Geoffrey Hinton · 2022
Later among the works it cites.
Diffupose: Monocular 3d human pose estimation via denoising diffusion probabilistic model
Jeongjun Choi, Dongseok Shim, and H Jin Kim · 2022
Later among the works it cites.
Diffusion models as plug-and-play priors
Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic, and Dimitris Samaras · 2022
Later among the works it cites.
Diffpose: Multi-hypothesis human pose estimation using diffusion models
Karl Holmquist and Bastian Wandt · 2022
Later among the works it cites.
Expectation-maximization contrastive learning for compact video-and-language representations
Peng Jin, Jinfa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David A Clifton, and Jie Chen · 2022
Later among the works it cites.
Restr: Convolution-free referring image segmentation using transformers
Namyup Kim, Dongwon Kim, Cuiling Lan, Wenjun Zeng, and Suha Kwak · 2022
Later among the works it cites.
Toward 3d spatial reasoning for human-like text-based visual question answering
Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen · 2022
Later among the works it cites.
Grounded language-image pre-training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Cris: Clip-driven referring image segmentation
Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu · 2022
Later among the works it cites.
Diffusion models for implicit image segmentation ensembles
Julia Wolleb, Robin Sandkühler, Florentin Bieder, Philippe Valmaggia, and Philippe C Cattin · 2022
Later among the works it cites.
Lavt: Language-aware vision transformer for referring image segmentation
Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr · 2022
Later among the works it cites.
Coupalign: Coupling word-pixel with sentence-mask alignments for referring image segmentation
Zicheng Zhang, Yi Zhu, Jianzhuang Liu, Xiaodan Liang, and Wei Ke · 2022
Later among the works it cites.
Seqtr: A simple yet universal network for visual grounding
Chaoyang Zhu, Yiyi Zhou, Yunhang Shen, Gen Luo, Xingjia Pan, Mingbao Lin, Chao Chen, Liujuan Cao, Xiaoshuai Sun, and Rongrong Ji · 2022
Later among the works it cites.
Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning
Jin Peng, Huang JinFa, Xiong Pengfei, Tian Shangxuan, Liu Chang, Ji Xiangyang, Yuan Li, and Chen Jie · 2023
Closest in time.