Fetching the paper…
Reading the bibliography…
In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-vocabulary scenario.
Describing objects by their attributes
Ali Farhadi, Ian Endres, Derek Hoiem, and David Forsyth · 2009
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling · 2009
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Attribute-based classification for zero-shot learning of object categories
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling · 2013
Earlier work this paper cites.
Selective search for object recognition
Jasper RR Uijlings, Koen EA Van De Sande, Theo Gevers, and Arnold WM Smeulders · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Zero-shot recognition with unreliable attributes
Dinesh Jayaraman and Kristen Grauman · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Multi-task cnn model for attribute prediction
Abrar H Abdulnabi, Gang Wang, Jiwen Lu, and Kui Jia · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph J Lim, and Edward H Adelson · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Coco attributes: Attributes for people, animals, and objects
Genevieve Patterson and James Hays · 2016
Earlier work this paper cites.
Face attribute prediction using off-the-shelf cnn features
Yang Zhong, Josephine Sullivan, and Haibo Li · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Zero-shot object detection
Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran · 2018
Earlier work this paper cites.
Zero-shot object detection: Learning to simultaneously recognize and localize novel concepts
Shafin Rahman, Salman Khan, and Fatih Porikli · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Deep imbalanced learning for face recognition and attribute prediction
Chen Huang, Yining Li, Chen Change Loy, and Xiaoou Tang · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Cited alongside, same era.
Task-aware attention model for clothing attribute prediction
Sanyi Zhang, Zhanjie Song, Xiaochun Cao, Hua Zhang, and Jie Zhou · 2019
Cited alongside, same era.
Zero-shot object detection with attributes-based category similarity
Qiaomei Mao, Chong Wang, Shenghao Yu, Ye Zheng, and Yuqi Li · 2020
Cited alongside, same era.
Adinet: Attribute driven incremental network for retinal image classification
Learning to predict visual attributes in the wild
Khoi Pham, Kushal Kafle, Zhe Lin, Zhihong Ding, Scott Cohen, Quan Tran, and Abhinav Shrivastava · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Open-vocabulary object detection using captions
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang · 2021
Later among the works it cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Yan Zeng, Xinsong Zhang, and Hang Li · 2021
Later among the works it cites.
Localized vision-language matching for open-vocabulary object detection
Maria A Bravo, Sudhanshu Mittal, and Thomas Brox · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qier Meng and Satoh Shin’ichi · 2020
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Cited alongside, same era.
Connecting vision and language with localized narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari · 2020
Cited alongside, same era.
Improved visual-semantic alignment for zero-shot object detection
Shafin Rahman, Salman Khan, and Nick Barnes · 2020
Cited alongside, same era.
Phrasecut: Language-based image segmentation in the wild
Chenyun Wu, Zhe Lin, Scott Cohen, Trung Bui, and Subhransu Maji · 2020
Cited alongside, same era.
Attribute prototype network for zero-shot learning
Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele, and Zeynep Akata · 2020
Cited alongside, same era.
Texture and shape biased two-stream networks for clothing classification and attribute recognition
Yuwei Zhang, Peng Zhang, Chun Yuan, and Zhi Wang · 2020
Cited alongside, same era.
Open-vocabulary attribute detection
María A Bravo, Sudhanshu Mittal, Simon Ging, and Thomas Brox · 2022
Later among the works it cites.
Promptdet: Expand your detector vocabulary with uncurated images
Chengjian Feng, Yujie Zhong, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, and Lin Ma · 2022
Later among the works it cites.
Label, verify, correct: A simple few shot object detection method
Prannay Kaul, Weidi Xie, and Andrew Zisserman · 2022
Later among the works it cites.
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Glidenet: Global, local and intrinsic based dense embedding network for multi-category attributes prediction
Kareem Metwaly, Aerin Kim, Elliot Branson, and Vishal Monga · 2022
Later among the works it cites.
Improving closed and open-vocabulary attribute prediction using transformers
Khoi Pham, Kushal Kafle, Zhe Lin, Zhihong Ding, Scott Cohen, Quan Tran, and Abhinav Shrivastava · 2022
Later among the works it cites.
Bridging the gap between object and image-level representations for open-vocabulary detection
Hanoona Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman Khan, and Fahad Shahbaz Khan · 2022
Later among the works it cites.
Disentangling visual embeddings for attributes and objects
Nirat Saini, Khoi Pham, and Abhinav Shrivastava · 2022
Later among the works it cites.
Attributes learning network for generalized zero-shot learning
Yu Yun, Sen Wang, Mingzhen Hou, and Quanxue Gao · 2022
Later among the works it cites.
Regionclip: Region-based language-image pretraining
Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krähenbühl, and Ishan Misra · 2022
Later among the works it cites.