Fetching the paper…
Reading the bibliography…
Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements.
“What makes a chair a chair?”
Helmut Grabner, Juergen Gall and Luc Van · 2011
Earlier work this paper cites.
“Referitgame: Referring to objects in photographs of natural scenes”
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten and Tamara Berg · 2014
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
“Affordance detection of tool parts from geometric features”
Austin Myers, Ching Teo, Cornelia Fermüller and Yiannis Aloimonos · 2015
Earlier work this paper cites.
“Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models”
Bryan Plummer et al · 2015
Earlier work this paper cites.
“Generation and Comprehension of Unambiguous Object Descriptions”
Junhua Mao et al · 2016
Earlier work this paper cites.
“Modeling Context in Referring Expressions”
Licheng Yu et al · 2016
Earlier work this paper cites.
“Modeling Relationships in Referential Expressions with Compositional Modular Networks”
Ronghang Hu et al · 2017
Earlier work this paper cites.
“Visual genome: Connecting language and vision using crowdsourced dense image annotations”
Ranjay Krishna et al · 2017
Earlier work this paper cites.
“Scene parsing through ade20k dataset”
Bolei Zhou et al · 2017
Earlier work this paper cites.
“Real-time Referring Expression Comprehension by Single-stage Grounding Network”
Xinpeng Chen et al · 2018
Earlier work this paper cites.
“Learning to act properly: Predicting and explaining affordances from images”
Ching-Yao Chuang, Jiaman Li, Antonio Torralba and Sanja Fidler · 2018
Earlier work this paper cites.
“Fixing Weight Decay Regularization in Adam”, 2018
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
“YOLOv3: An Incremental Improvement”
Joseph Redmon and Ali Farhadi · 2018
Earlier work this paper cites.
“Learning Two-Branch Neural Networks for Image-Text Matching Tasks”
Liwei Wang, Yin Li, Jing Huang and Svetlana Lazebnik · 2018
Earlier work this paper cites.
“MAttNet: Modular Attention Network for Referring Expression Comprehension”
Licheng Yu et al · 2018
Earlier work this paper cites.
“Grounding Referring Expressions in Images by Variational Context”
Hanwang Zhang, Yulei Niu and Shih-Fu Chang · 2018
Earlier work this paper cites.
“Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries”
Bohan Zhuang et al · 2018
Cited alongside, same era.
“See-through-text grouping for referring image segmentation”
Ding-Jie Chen et al · 2019
Cited alongside, same era.
“UNITER: Learning Universal Image-Text Representations”, 2019
Yen-Chun Chen et al · 2019
Cited alongside, same era.
“Learning to Compose and Reason with Language Tree Structures for Visual Grounding”
Richang Hong et al · 2019
Cited alongside, same era.
“Learning to Assemble Neural Module Tree Networks for Visual Grounding”
Daqing Liu, Hanwang Zhang, Feng Wu and Zheng-Jun Zha · 2019
Cited alongside, same era.
“A Real-time Cross-modality Correlation Filtering Method for Referring Expression Comprehension”
Yue Liao et al · 2020
Later among the works it cites.
“12-in-1: Multi-task vision and language representation learning”
Jiasen Lu et al · 2020
Later among the works it cites.
“Cascade grouped attention network for referring expression segmentation”
Gen Luo et al · 2020
Later among the works it cites.
“VL-BERT: Pre-training of Generic Visual-Linguistic Representations”
Weijie Su et al · 2020
Later among the works it cites.
“Improving One-Stage Visual Grounding by Recursive Sub-Query Construction”
Zhengyuan Yang, Tianlang Chen, Liwei Wang and Jiebo Luo · 2020
Later among the works it cites.
“Vision-language transformer and query generation for referring segmentation”
Henghui Ding, Chang Liu, Suchen Wang and Xudong Jiang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu et al · 2019
Cited alongside, same era.
“ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks”
Jiasen Lu, Dhruv Batra, Devi Parikh and Stefan Lee · 2019
Cited alongside, same era.
“Zero-Shot Grounding of Objects from Natural Language Queries”
Arka Sadhu, Kan Chen and Ram Nevatia · 2019
Cited alongside, same era.
“What object should i use?-task driven object detection”
Johann Sawatzky, Yaser Souri, Christian Grund and Jurgen Gall · 2019
Cited alongside, same era.
“Neighbourhood Watch: Referring Expression Comprehension via Language-Guided Graph Attention Aetworks”
Peng Wang et al · 2019
Cited alongside, same era.
“Dynamic Graph Attention for Referring Expression Comprehension”
Sibei Yang, Guanbin Li and Yizhou Yu · 2019
Cited alongside, same era.
“A Fast and Accurate One-Stage Approach to Visual Grounding”
Zhengyuan Yang et al · 2019
Cited alongside, same era.
Later among the works it cites.
“Encoder fusion network with co-attention embedding for referring image segmentation”
Guang Feng, Zhiwei Hu, Lihe Zhang and Huchuan Lu · 2021
Later among the works it cites.
“MDETR-Modulated Detection for End-to-End Multi-Modal Understanding”
Aishwarya Kamath et al · 2021
Later among the works it cites.
“Referring transformer: A one-step approach to multi-task visual grounding”
Muchen Li and Leonid Sigal · 2021
Later among the works it cites.
“One-shot affordance detection”
Hongchen Luo et al · 2021
Later among the works it cites.
“Phrase-based affordance detection via cyclic bilateral interaction”
Liangsheng Lu et al · 2022
Later among the works it cites.
“SiRi: A Simple Selective Retraining Mechanism for Transformer-Based Visual Grounding”
Mengxue Qu et al · 2022
Later among the works it cites.
“One-shot object affordance detection in the wild”
Wei Zhai et al · 2022
Later among the works it cites.
“Seqtr: A simple yet universal network for visual grounding”
Chaoyang Zhu et al · 2022
Later among the works it cites.
“Polyformer: Referring image segmentation as sequential polygon generation”
Jiang Liu et al · 2023
Closest in time.
“Grounding dino: Marrying dino with grounded pre-training for open-set object detection”
Shilong Liu et al · 2023
Closest in time.